KEP-5517: DRA Node Allocatable Resources
KEP-5517: DRA: Node Allocatable Resources
- Release Signoff Checklist
- Summary
- Motivation
- Proposal
- Design Details
- Conceptual Mapping: Pod Spec Requests and Limits with DRA
- API Changes
- Kube-Scheduler Changes
- Node Resource Enforcement and Isolation
- Kubelet Admission Control
- ResourceQuota Enforcement
- HPA Integration
- Cluster Autoscaler Integration
- Node Capacity Reporting
- Future Enhancements
- Test Plan
- Graduation Criteria
- Upgrade / Downgrade Strategy
- Version Skew Strategy
- Production Readiness Review Questionnaire
- Implementation History
- Drawbacks
- Alternatives
- Infrastructure Needed (Optional)
Release Signoff Checklist
Items marked with (R) are required prior to targeting to a milestone / release.
- (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
- (R) KEP approvers have approved the KEP status as
implementable - (R) Design details are appropriately documented
- (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
- e2e Tests for all Beta API Operations (endpoints)
- (R) Ensure GA e2e tests meet requirements for Conformance Tests
- (R) Minimum Two Week Window for GA e2e tests to prove flake free
- (R) Graduation criteria is in place
- (R) all GA Endpoints must be hit by Conformance Tests within one minor version of promotion to GA
- (R) Production readiness review completed
- (R) Production readiness review approved
- “Implementation History” section is up-to-date for milestone
- User-facing documentation has been created in kubernetes/website, for publication to kubernetes.io
- Supporting documentation—e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes
Summary
This KEP proposes a solution for managing node allocatable resources via Dynamic Resource Allocation (DRA). Node allocatable resources are resources currently reported in v1.Node status.allocatable that are not extended resources (examples include CPU, Memory, Ephemeral-storage, and Hugepages). Currently, when these node allocatable resources are managed via DRA, there is a fundamental disconnect across the control plane and the Node. In the scheduler, having two independent accounting systems (one for standard resources, one for DRA) managing the same underlying resource leads to resource overcommitment. On the node, the kubelet is completely unaware of DRA allocations, which may result in incorrect QoS class assignment and has many downstream implications. This forces users into fragile workarounds that are incompatible with all use cases.
The proposed solution in this KEP addresses node allocatable resource accounting and enforcement in kube-scheduler and kubelet:
- Kube-Scheduler Accounting: During filtering and scoring,
NodeResourcesFitandDynamicResourcesaccount for node allocatable resources allocated via DRAResourceClaims both for the incoming pod being scheduled and for pods already running on the node preventing node overcommitment. - Kubelet Enforcement: kubelet natively incorporates node allocatable resource allocations made through DRA
ResourceClaims to configure Linux container and pod cgroups and calculate OOM score.
Motivation
Dynamic Resource Allocation (DRA) provides a powerful framework for managing specialized hardware resources such as GPUs, FPGAs, and high-performance network interfaces. It also enables fine-grained management of node allocatable resources like CPU and Memory, for example, through the dra-driver-cpu. However, when a node allocatable resource is managed via DRA, while it provides added advantages of being able to specify more detailed requirements, a fundamental disconnect emerges between the scheduler, the kubelet, and the DRA framework, which breaks the resource guarantees.
Additionally, specialized resources like accelerators often have implicit dependencies on node allocatable resources like CPU or Hugepages for the application to interact with it. Currently, users must manually research and declare these auxiliary node allocatable resource requirements, typically as additional requests in the PodSpec. This process is error-prone and adds complexity to workload configuration. Furthermore, there is no existing mechanism to express critical co-location requirements. For example, there is no way to ensure an accelerator allocated via DRA is NUMA-aligned with the specific hugepages or CPUs it needs, as the standard and DRA resource models are entirely independent.
Core Problem
The core problem is that the same underlying physical resource is advertised and consumed through two parallel, uncoordinated mechanisms.
Dual Publication: A node’s total CPU/Memory capacity is advertised in two different places:
- Via the Kubelet in the
Node.Status.Allocatablefield. - Via the DRA driver in
ResourceSliceobjects.
- Via the Kubelet in the
Dual Consumption: Pods can consume this CPU capacity in two different ways:
- Via pod spec requests (
pod.spec.containers[].resources.requests,pod.spec.initcontainers[].resources.requests), which is considered in theNodeResourcesFitscheduler plugin to find a Node that fits. - Via
ResourceClaim, which is considered in theDynamicResourcesscheduler plugin to allocate devices.
- Via pod spec requests (
Scheduler-Level Resource Oversubscription: The kubelet is the source of truth for a node’s
available resources. The scheduler continuously watches the Node object and uses
Node.Status.Allocatable to maintain an internal, in-memory cache (NodeInfo) of each node’s
capacity. This cache is the baseline for all its scheduling decisions, ensuring it does not place
more pods on a node than the node reports it can handle.
It is completely blind to the fact that the DRA (like CPU ResourceClaim) draws from the same
physical resource as a standard request. This gap leads to the scheduler overcommitting a node’s CPU
resources by scheduling more pods than the node resource capacity.
Kubelet-Level Guarantee Failure: The kubelet is the component that enforces resource guarantees on the node. It configures Linux cgroups, calculates Out-Of-Memory (OOM) score adjustments, and makes critical lifecycle decisions like eviction based only on standard pod.Spec requests and limits. Because Kubelet is unaware of resources allocated via DRA, workloads suffer from an Enforcement Gap:
- Even if the scheduler correctly reserves capacity for both standard and DRA requests on a node, the container remains hard-restricted by the Kubelet’s Linux cgroups to its standard Spec bounds. For example, if a container requests 2 CPU in its Spec and references a claim for 5 CPU, the container runtime applies a cgroup CPU quota of only 2 CPU. If the application attempts to consume the 5 CPU burst allocated via DRA, it will be hard-throttled by the kernel.
- If a workload relies on memory provided via a DRA claim but its standard Spec memory limit is lower:
- The kernel will terminate the container when its usage exceeds the standard memory limit.
- Kubelet sets a higher OOM score based strictly on the smaller standard memory request, making the workload a prime target for the kernel OOM kill during host memory exhaustion.
Current workarounds for DRA-managed node allocatable resources (like
CPU DRA driver) force users to duplicate
resource requests in both the ResourceClaim and the standard pod.spec.containers[].resources.
However, this approach is fragile, error-prone, and difficult to manage, especially for complex pods
with shared resource claims. It is also incompatible with advanced DRA features like
Prioritized Lists
This KEP proposes to solve this problem by creating a single, unified resource model that spans the entire control plane, from the scheduler to the kubelet. The goal is not just to fix an accounting issue in the scheduler, but to provide a complete, native way for Kubernetes to handle core resources that are backed by DRA.
Goals
- To create a unified accounting model within the kube-scheduler that prevents overcommitment of core
resources (like CPU) when they are allocated via both standard
pod.specrequests and DRAResourceClaims. - To ensure the solution is compatible with different ways node allocatable resources can be represented and allocated within DRA, including as individual devices, consumable capacities (KEP-5075), and partitionable devices (KEP-4815)
- To enable specialized devices, such as accelerators, to declare any auxiliary node allocatable resource requirements (e.g., CPU, Memory) they depend on for their operation.
- To natively integrate DRA node allocatable resource allocations into Kubelet cgroup enforcement.
- To maintain backward compatibility with existing workloads and ecosystem tools that rely on
node.status.allocatableand the scheduler’s view of node resource utilization.
Non-Goals
- To move all resource management logic into the DRA driver. The Kubelet will remain the primary agent for cgroup management and QoS enforcement, ensuring that the benefits of its existing stability and lifecycle management features are preserved.
- To replace the standard
pod.spec.containers.resourcesAPI for requesting node allocatable resources. This KEP aims to enhance the system by adding a clear path for node allocatable resource requests via DRA while ensuring it works coherently with the existing PodSpec-based requests. - Modifying Kubelet’s core QoS class classification logic is a non-goal for this KEP. QoS will still be based strictly on standard Spec requests and limits.
Proposal
This KEP introduces a unified accounting and enforcement model within kube-scheduler and the Kubelet to integrate
node allocatable resources managed by Dynamic Resource Allocation (DRA) with standard resource tracking. By bridging
the gap between pod.spec.resources and DRA ResourceClaim allocations, we can achieve consistent resource
accounting and prevent node overcommitment.
Background
To understand the proposed solution, it is essential to first understand how the control plane and the node currently manage standard resource requests and DRA ResourceClaims.
Kube-Scheduler Background
The Kubernetes scheduler is built on a plugin-based framework that executes a series of stages to place
a pod. This KEP is primarily concerned with the interaction between NodeResourcesFit and
DynamicResources plugins across the PreFilter, Filter, Score, PreBind, and Bind stages of the
scheduling framework.
Standard Resource Accounting
The Kubelet is the source of truth for a node’s available resources. It inspects the machine’s total
capacity, subtracts resources reserved for the operating system (--system-reserved) and Kubernetes
system daemons (--kube-reserved), and reports the result in the Node.Status.Allocatable field. The
scheduler continuously watches for updates to this field and uses it to maintain its internal, in-memory
cache (NodeInfo) of each node’s capacity. This cache is the baseline for all its scheduling decisions.
Kube-Scheduler Resource Accounting
- The scheduler maintains an in-memory
NodeInfoobject for each node, which stores theAllocatable, which is the capacity of the node andRequested, which is an aggregated sum of the resources requested by all pods assumed to be on that node (Requested). - During the
Filterstage of scheduling, theNodeResourcesFitplugin checks if a pod’s requested resources can fit on the node (NodeInfo.Allocatable - NodeInfo.Requested >= Pod request). - The
NodeInfo.Requestedvalue is updated by the scheduler framework when a pod is “assumed” on the node. This happens after a node is selected in theScoringphase, and before the actual binding to the API server, ensuring the cache is accurate for subsequent scheduling decisions.
Dynamic Resource Allocation (DRA) Accounting
The DynamicResources plugin manages resources requested via pod.spec.resourceClaims. Its accounting
system is entirely separate from the standard resources.
- The DRA driver/s on the node reports resource availability through the
ResourceSliceobjects. - During the
Filterstage, theDynamicResourcesplugin determines if the inventory in theResourceSliceobjects is sufficient to satisfy the pod’sResourceClaim, after accounting for devices already allocated to other claims. - When a pod is scheduled, the
DynamicResourcesplugin, in itsPreBindstage, makes an API call to update theResourceClaimobject’s status. This update makes the allocation permanent and visible to the rest of the cluster.
The standard resource and dynamic resource accounting systems are completely independent. The
NodeInfo cache is not aware of allocations recorded in ResourceClaim objects, which is the root
cause of the accounting gap for node allocatable resources when they are managed through DRA.
Node Resource Enforcement Background
To enforce physical resource guarantees and isolation on the host, the Kubelet configures the kernel cgroup settings and Out-Of-Memory (OOM) score adjustments based on the pod specification.
Cgroup Enforcement
The Kubelet establishes resource boundaries at both the top-level pod cgroup and individual container cgroups via the Container Runtime Interface (CRI):
- Container-Level cgroups: By default, the Kubelet translates the requests and limits specified in
pod.Spec.Containers[].Resourcesdirectly into container-level cgroup parameters:- CPU Requests establish the relative weight (
cpu.weightorcpu.shares) for fair scheduling during machine contention. - CPU Limits configure the hard threshold (
cpu.maxorcpu.cfs_quota_us). Workloads attempting to burst above this threshold are throttled by the kernel. - Memory Limits set the memory usage threshold (
memory.maxormemory.limit_in_bytes). Exceeding this limit triggers an immediate Out-Of-Memory kill.
- CPU Requests establish the relative weight (
- Pod-Level cgroups: When Pod Level Resources (
pod.spec.resources) are explicitly specified, the Kubelet applies the overall resource request and limit directly to the parent pod-level cgroup.- The aggregate resource consumption of all containers combined (including init, sidecar, and regular containers) is hard-capped by this pod-level limit.
- If an individual container omits its own limit while a pod-level limit is set, the Kubelet applies the pod-level limit to that container’s cgroup maximum value. This explicit fallback is critical because container-level limits are implied under a pod budget, and runtimes (such as the Java Virtual Machine) inspect container-level cgroup maximums to fine-tune internal memory pools and thread allocations.
- If pod level resources are not explicitly specified, the Kubelet sums up the container-level resource requests and limits and sets pod-level cgroups
OOM Score Adjustments
To ensure node stability during memory exhaustion, the Kubelet configures the oom_score_adj parameter for each container. This value informs the Linux kernel OOM killer which processes to terminate first:
- For Guaranteed and BestEffort pods, the Kubelet applies static constant scores (
-997and1000). - For Burstable pods, the score is dynamically calculated based on the container’s standard memory requests relative to the node’s memory capacity. Higher memory requests yield more protective (lower) scores, reducing the likelihood of premature termination.
User Stories
Story 1 (Resource Alignment): An HPC workload needs a certain number of exclusive CPUs and memory
that are aligned on the same NUMA node as a specific NIC for maximum performance. The user creates a
ResourceClaim with co-location constraints to enforce this. The scheduler correctly accounts for the
CPU and memory requests made through the claim, adding them to the node’s total requested resources, so
the node is not oversubscribed.
Story 2 (Dedicated and Shared resources): A telco application has some high-priority application
containers and some lower-priority sidecar containers. The user wants to dedicate some CPU cores
exclusively to the application containers for low latency, while allowing sidecar containers to run on
the node’s general shared CPU pool. They use DRA to request exclusive cores and standard pod.spec
requests for the shared CPU portion. The scheduler should correctly account for both dedicated and shared
requests made through these different mechanisms.
Story 3 (Accelerator with Node Allocatable Resource Dependency): An AI inference job requests a GPU through
a ResourceClaim. The specific GPU model also requires a certain number of CPUs and Hugepages that are
required for the application to interact with the accelerator. Instead of requiring the user to know
about these auxiliary CPU and HugePages requests and add it to their PodSpec, the GPU device can be configured to declare these dependencies. The Kubernetes scheduler accounts for both the CPU/HugePages
needs for the GPU device and the standard pod spec requests, ensuring the pod lands on a node with
sufficient capacity for all requirements. The user experience is simplified, as they only need to ask
for the primary device they care about.
Story 4 (Fungibility): An ML inference job can use either a full GPU or, if none is available, a
slice of 8 exclusive CPUs. The user creates a ResourceClaim with a firstAvailable list to
represent this fungible need. The scheduler evaluates both paths against a node’s available
resources. It finds a node with 8 available CPUs, correctly reserves them in its central NodeInfo
cache, and schedules the pod. The user did not need to guess which resource to put in the pod.spec.
Risks and Mitigations
- Increased API and user complexity by having two ways to request node allocatable resources (PodSpec and ResourceClaim). To mitigate, the documentation would be enhanced with clear guidelines and use cases for DRA for Node Allocatable Resources.
- Bugs in the kube-scheduler’s new accounting logic could lead to incorrect node resource calculations and node oversubscription. Extensive unit and integration tests covering various resource claim and standard request combinations should help mitigate this. The feature will also be rolled out gradually, beginning with an alpha release to gather feedback and address potential concerns.
- While the Kubelet considers DRA for cgroup enforcement, QoS class classification remains purely based on the standard Spec.
Pods that only use DRA claims to request node allocatable resources are classified as
BestEffortpods and are more susceptible to node eviction and stricter cgroup enforcement compared to pods requesting the same amount of resources through standard requests. This is discussed in the QoS Class Mismatch Risks section.
Design Details
The proposal here is to implement a “Unified Accounting and Enforcement” model across the control plane and the host for node allocatable resources requested through the standard pod Spec or through Dynamic Resource Allocation (DRA) claims. This involves:
- API Changes: Updates to the DRA API for drivers to declare node allocatable resource implications in
Deviceobjects, and PodStatus to record DRA-based node allocatable resource allocations. - Kube-Scheduler Changes: Both
NodeResourcesFitandDynamicResourcesevaluate the pod’s accumulated resource requirement (i.e., standard requests plus resolved DRA claim allocations) so the fit check is independent of plugin ordering. Resource scoring accounts for the DRA footprint of both existing pods on the node and the candidate pod being scored. - Kubelet Changes: Updates in Kubelet to take into account resources allocated through DRA in the cgroup enforcement.
- Kube-Controller-Manager Changes: Updates to Resource Quota evaluation and the Horizontal Pod Autoscaler (HPA) to account for DRA node allocatable resources recorded in
PodStatus.
Conceptual Mapping: Pod Spec Requests and Limits with DRA
Traditional resources like CPU and Memory in the Pod Spec have allocations split into requests (for capacity reservation and cgroup weight) and limits (for hard cgroup ceilings). Since DRA is primarily used for hardware devices like accelerators and NICs, DRA API lacks the concept of separate requests and limits. To bridge standard resource enforcements with DRA claims, we use DRA allocations along with traditional requests and limits as follows:
- In the Scheduler: In addition to standard requests, the DRA allocation acts as a request to deduct capacity from the node and prevent overcommitment.
- On the Node: The DRA allocation acts as both a request (cgroup shares/weight) to enforce pod level cgroup bounds based on the scheduler-reserved resource footprint and a limit to allow the containers to utilize the capacity.
Importantly, these DRA allocations are strictly additive to the standard resources declared in the Pod Spec; they enhance cgroup boundaries without replacing the existing Pod Spec-based requests and limits.
API Changes
To support unified accounting for node allocatable resources, this KEP proposes API extensions to the Device object and PodStatus.
Device API Extensions
The new field NodeAllocatableResources within the
ResourceSlice.Device spec is used to define the node allocatable
resource quantities.
// In k8s.io/api/resource/v1/types.go
type Device struct {
// ... existing fields
// NodeAllocatableResources defines the mapping of node resources
// that are managed by the DRA driver exposing this device. These are resources currently
// reported in v1.Node `status.allocatable` that are not extended resources
// (see https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/#extended-resources).
// The only allowed keys are "cpu", "memory", "hugepages-<size>", and "ephemeral-storage".
// In addition to standard requests made through the Pod `spec`, these resources
// can also be requested through claims and allocated by the DRA driver.
// For example, a CPU DRA driver might allocate exclusive CPUs or auxiliary node memory
// dependencies of an accelerator device.
// The keys of this map are the node allocatable resource names (e.g., "cpu", "memory", "ephemeral-storage").
// Extended resource names are not permitted as keys.
// +optional
// +featureGate=DRANodeAllocatableResources
NodeAllocatableResources map[v1.ResourceName]NodeAllocatableResource `json:"nodeAllocatableResources,omitempty" protobuf:"bytes,14,opt,name=nodeAllocatableResources"`
}
// NodeAllocatableResource defines the translation between the DRA device/capacity
// units requested to the corresponding quantity of the node allocatable resource.
// At least one of Mapping or Overhead must be specified. Not specifying either is an invalid configuration.
type NodeAllocatableResource struct {
// Mapping is used when the device directly models a node allocatable resource like standard CPU or memory
// (e.g., with a CPU DRA driver). The calculated quantity is accounted for exactly once per claim instance
// on the node. To prevent node cgroup isolation friction, the scheduler explicitly
// blocks sharing mapped device claims across multiple pods.
// +optional
// +k8s:optional
Mapping *NodeAllocatableMapping `json:"mapping,omitempty" protobuf:"bytes,3,opt,name=mapping"`
// Overhead contains fields for modeling auxiliary overhead incurred on node allocatable resources
// when allocating devices that are not themselves modeling a node allocatable resource (e.g., host memory overhead for GPUs).
// Sharing overhead-mapped claims across multiple pods is allowed. The node allocatable overhead is accounted
// for individually for each pod referencing the claim.
// Overhead is always subtracted from the node's allocatable capacity for the resource, even when mapping
// is specified for the same resource.
// Eg: If a device models memory capacity per socket as a consumable capacity pool via Mapping (with CapacityKey),
// any overhead specified for the same resource will be subtracted from the node's general allocatable capacity
// and not from the per-socket capacity pool in Mapping.
// +optional
// +k8s:optional
Overhead *NodeAllocatableOverhead `json:"overhead,omitempty" protobuf:"bytes,4,opt,name=overhead"`
}
// NodeAllocatableMapping defines how a DRA allocation directly translates into a node allocatable resource quantity.
// The mapping can be derived from either the count of allocated devices (via deviceMultiplier) or the specific capacity consumed (via capacityKey and capacityMultiplier). These options are mutually exclusive.
// Kubelet adds this mapped resource quantity from claim to both requests and limits at the pod-level cgroup, and to limits at the container-level cgroup for each container referencing the claim.
type NodeAllocatableMapping struct {
// CapacityKey references a capacity name defined as a key in the
// `spec.devices[*].capacity` map. When this field is set, the value associated with
// this key in the `status.allocation.devices.results[*].consumedCapacity` map
// (for a specific claim allocation) determines the base quantity for
// the node allocatable resource. `capacityMultiplier` must also be set and is
// multiplied with the base quantity.
// For example, if `spec.devices[*].capacity` has an entry "dra.example.com/memory": "128Gi",
// and this field is set to "dra.example.com/memory", then for a claim allocation
// that consumes { "dra.example.com/memory": "4Gi" } the base quantity for the
// node allocatable resource mapping will be "4Gi".
// The final node allocatable resource amount is `consumedCapacity[capacityKey]` * `capacityMultiplier`.
// +optional
// +k8s:optional
// +k8s:unionMember
// +k8s:alpha(since: "1.37")=+k8s:dependentRequired("capacityMultiplier")
CapacityKey *QualifiedName `json:"capacityKey,omitempty" protobuf:"bytes,1,opt,name=capacityKey"`
// CapacityMultiplier is used as a multiplier for the allocated capacity consumed.
// It is only valid if `capacityKey` is set.
// The final node allocatable resource amount is `consumedCapacity[capacityKey]` * `capacityMultiplier`.
// For example, if a Device's capacity "dra.example.com/cores" is consumed,
// and each "core" provides 2 "cpu"s, the mapping would be:
// {ResourceName: "cpu", capacityKey: "dra.example.com/cores", capacityMultiplier: "2"}.
// If a claim consumes 8 "dra.example.com/cores", the CPU footprint is 8 * 2 = 16.
// +optional
// +k8s:optional
// +k8s:alpha(since: "1.37")=+k8s:dependentRequired("capacityKey")
CapacityMultiplier *resource.Quantity `json:"capacityMultiplier,omitempty" protobuf:"bytes,2,opt,name=capacityMultiplier"`
// DeviceMultiplier is used as a multiplier for the allocated device count in the claim.
// The final node allocatable resource amount is `deviceCount` * `deviceMultiplier`.
// For example, a DRA driver representing each cache complex (CCX) as a device would have
// {ResourceName: "cpu", deviceMultiplier: "8"} in its `nodeAllocatableResources`.
// If 2 devices (CCX) are allocated to the claim, 2 * 8 = 16 CPUs would be considered as allocated.
// It is only valid when `capacityKey` and `capacityMultiplier` are not set.
// +optional
// +k8s:optional
// +k8s:unionMember
DeviceMultiplier *resource.Quantity `json:"deviceMultiplier,omitempty" protobuf:"bytes,3,opt,name=deviceMultiplier"`
}
// NodeAllocatableOverhead defines auxiliary resource overheads incurred when allocating a device.
// Overheads can be specified as a fixed cost per pod referencing the claim, a variable cost per container reference, or both.
// Kubelet accounts for this overhead by adding it to both the pod-level and container-level cgroups of referencing containers.
type NodeAllocatableOverhead struct {
// PerPod is overhead applied once per pod referencing the claim on this node.
// This is a flat overhead incurred for every pod referencing the claim.
// +optional
// +k8s:optional
PerPod *resource.Quantity `json:"perPod,omitempty" protobuf:"bytes,1,opt,name=perPod"`
// PerContainer is applied per container reference to the claim.
// This models overhead scaling linearly with the number of containers actively using the device.
// When both PerPod and PerContainer are specified, the total overhead allocated for each pod referencing
// the claim is computed as:
// Quantity = PerPod + (PerContainer * NumReferences)
// Kubelet accounts for this overhead in cgroups:
// - Pod-level cgroup (requests and limits): Kubelet adds PerPod + (PerContainer * NumReferences).
// - Container-level cgroup (limits only): Kubelet adds PerPod + PerContainer for each referencing container.
// This allows any single container to access the pod-level overhead, while the parent cgroup caps the total usage to account for PerPod exactly once.
// +optional
// +k8s:optional
PerContainer *resource.Quantity `json:"perContainer,omitempty" protobuf:"bytes,2,opt,name=perContainer"`
}
Pod API Changes
We add a new field AdditionalNodeAllocatableResources to PodStatus (renamed from NodeAllocatableResourceClaimStatuses in Alpha) as a way to pass the allocation details from the DynamicResources plugin to the kube-scheduler accounting logic.
// In k8s.io/api/core/v1/types.go
// PodStatus represents information about the status of a pod.
type PodStatus struct {
// ... existing fields
// NodeAllocatableResourceClaimStatuses is tombstoned since it got replaced with AdditionalNodeAllocatableResources.
// NodeAllocatableResourceClaimStatuses []NodeAllocatableResourceClaimStatus `json:"nodeAllocatableResourceClaimStatuses,omitempty" patchStrategy:"merge" patchMergeKey:"resourceClaimName" protobuf:"bytes,21,rep,name=nodeAllocatableResourceClaimStatuses"`
// AdditionalNodeAllocatableResources contains the status of node allocatable resources
// that were allocated for this pod outside of direct spec requests (e.g., through DRA claims).
// This includes resources currently
// reported in v1.Node `status.allocatable` that are not extended resources
// (see https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/#extended-resources).
// Examples include "cpu", "memory", "ephemeral-storage", and hugepages.
// +featureGate=DRANodeAllocatableResources
// +optional
// +listType=atomic
AdditionalNodeAllocatableResources []AdditionalNodeAllocatableResource `json:"additionalNodeAllocatableResources,omitempty" protobuf:"bytes,25,rep,name=additionalNodeAllocatableResources"`
}
// AdditionalNodeAllocatableResource describes the status of
// node allocatable resources allocated outside of direct spec requests.
type AdditionalNodeAllocatableResource struct {
// ResourceClaimName is tombstoned since it got replaced with Source.
// ResourceClaimName string `json:"resourceClaimName" protobuf:"bytes,1,opt,name=resourceClaimName"`
// Source identifies the object in the pod's namespace that this resource
// contribution originates from (e.g., a ResourceClaim).
// +required
// +k8s:required
Source AdditionalNodeAllocatableReference `json:"source" protobuf:"bytes,6,opt,name=source"`
// Containers lists the names of all containers in this pod that reference the source.
// +optional
// +listType=set
// +k8s:optional
// +k8s:listType=set
Containers []string `json:"containers,omitempty" protobuf:"bytes,2,rep,name=containers"`
// Resources is tombstoned since it got replaced with more granular Mapping and Overhead fields.
// Resources map[ResourceName]resource.Quantity `json:"resources,omitempty" protobuf:"bytes,3,rep,name=resources"`
// Mapping contains fixed node allocatable resource quantities allocated once per source.
// When source.kind is ResourceClaim, this contains allocations through devices mapped in the device spec's `nodeAllocatableResources[...].mapping` field.
// This is used by kubelet for pod level and container-level cgroup enforcement.
// +optional
// +patchStrategy=merge
// +patchMergeKey=name
// +listType=map
// +listMapKey=name
// +k8s:optional
// +k8s:listType=map
// +k8s:listMapKey=name
Mapping []NodeAllocatableMappedResources `json:"mapping,omitempty" patchStrategy:"merge" patchMergeKey:"name" protobuf:"bytes,4,rep,name=mapping"`
// Overhead contains variable node allocatable resource overheads incurred per pod (PerPod) and
// per referencing container (PerContainer) when using the source.
// When source.kind is ResourceClaim, this contains allocations through devices mapped in the device spec's `nodeAllocatableResources[...].overhead` field.
// This is used by kubelet for pod level and container-level cgroup enforcement.
// +optional
// +patchStrategy=merge
// +patchMergeKey=name
// +listType=map
// +listMapKey=name
// +k8s:optional
// +k8s:listType=map
// +k8s:listMapKey=name
Overhead []NodeAllocatableOverheadResources `json:"overhead,omitempty" patchStrategy:"merge" patchMergeKey:"name" protobuf:"bytes,5,rep,name=overhead"`
}
// AdditionalNodeAllocatableReference identifies the source object in the pod's namespace
// contributing additional node allocatable resources to a pod.
// +structType=atomic
type AdditionalNodeAllocatableReference struct {
// APIGroup is the group for the resource being referenced. It is
// empty for the core API.
// Currently, only "resource.k8s.io" is supported.
// +optional
APIGroup string `json:"apiGroup,omitempty" protobuf:"bytes,1,opt,name=apiGroup"`
// Kind is the type of resource being referenced.
// Currently, only "ResourceClaim" is supported.
// +required
// +k8s:required
Kind string `json:"kind" protobuf:"bytes,2,opt,name=kind"`
// Name is the name of the source object in the pod's namespace
// (e.g., name of the resource claim).
// +required
// +k8s:required
Name string `json:"name" protobuf:"bytes,3,opt,name=name"`
}
// NodeAllocatableMappedResources describes mapped node allocatable resource allocations.
type NodeAllocatableMappedResources struct {
// Name is the name of the resource (e.g., cpu, memory).
// +required
// +k8s:required
Name ResourceName `json:"name" protobuf:"bytes,1,opt,name=name,casttype=ResourceName"`
// Quantity is the total node allocatable resource capacity allocated for the source.
// This source's allocated devices are shared by all the containers referencing the source.
// Kubelet adds this value to both requests and limits at the pod-level cgroup, and to limits at the container-level cgroup for each container referencing the source.
// +required
// +k8s:required
Quantity *resource.Quantity `json:"quantity" protobuf:"bytes,2,opt,name=quantity"`
}
// NodeAllocatableOverheadResources describes auxiliary overhead resource allocations.
type NodeAllocatableOverheadResources struct {
// Name is the name of the resource (e.g., cpu, memory).
// +required
// +k8s:required
Name ResourceName `json:"name" protobuf:"bytes,1,opt,name=name,casttype=ResourceName"`
// PerPod is the flat overhead quantity allocated per pod.
// Adding to each container limit allows individual containers to utilize the overhead, while the parent pod-level cgroup limit caps the total usage at the pod boundary where the overhead is accounted for exactly once.
// At least one of PerPod or PerContainer must be specified. Specifying neither is an invalid configuration.
// +optional
// +k8s:optional
PerPod *resource.Quantity `json:"perPod,omitempty" protobuf:"bytes,2,opt,name=perPod"`
// PerContainer is the variable overhead quantity applied for each container referencing the source.
// The container references are recorded in `additionalNodeAllocatableResources.containers`.
// The total overhead quantity allocated for the source is computed as:
// Quantity = PerPod + (PerContainer * NumReferences)
// Kubelet accounts for this overhead in cgroups:
// - Pod-level cgroup (requests and limits): Kubelet adds PerPod + (PerContainer * NumReferences).
// - Container-level cgroup (limits only): Kubelet adds PerPod + PerContainer for each referencing container.
// This allows any single container to access the pod-level overhead, while the parent cgroup caps the total usage to account for PerPod exactly once.
// At least one of PerPod or PerContainer must be specified. Specifying neither is an invalid configuration.
// +optional
// +k8s:optional
PerContainer *resource.Quantity `json:"perContainer,omitempty" protobuf:"bytes,3,opt,name=perContainer"`
}
Resource Representation Examples
- Direct Device Mapping with Individual Devices
- Each device instance in the slice corresponds directly to a fixed unit of the node allocatable resource.
- The
deviceMultiplierdetermines the resource footprint per device instance. - The number of devices allocated to the claim multiplied by
deviceMultiplierdetermines the overall node allocatable resource footprint and is recorded in the pod status.
# ResourceSlice
apiVersion: resource.k8s.io/v1
kind: ResourceSlice
metadata:
name: cpu-slice
spec:
driver: dra.example.com
nodeName: my-node
pool: { name: "node-pool", generation: 1, resourceSliceCount: 1 }
devices:
- name: cpu0
attributes: { numaNode: 0 }
nodeAllocatableResources:
cpu:
mapping:
deviceMultiplier: "1"
- name: cpu1
attributes: { numaNode: 0 }
nodeAllocatableResources:
cpu:
mapping:
deviceMultiplier: "1"
---
# ResourceClaim
apiVersion: resource.k8s.io/v1
kind: ResourceClaim
metadata:
name: cpu-claim
spec:
devices:
requests:
- name: cpu-req
exactly:
deviceClassName: cpu-core
count: 2
---
# Pod
apiVersion: v1
kind: Pod
metadata:
name: pod1
spec:
containers:
- name: worker
resources:
claims:
- name: my-cpu-claim
resourceClaims:
- name: my-cpu-claim
resourceClaimName: cpu-claim
status:
additionalNodeAllocatableResources:
- source:
apiGroup: resource.k8s.io
kind: ResourceClaim
name: cpu-claim
containers:
- worker
mapping:
- name: cpu
quantity: "2" # Derived from 2 allocated devices * multiplier 1
- Direct Device Mapping with Consumable Capacity
- The device is represented as a consumable capacity.
- The
capacityKeylinks the mapping directly to a specific capacity attribute inside the device. - The scheduler reads the exact consumed capacity from the claim allocation results to determine the base quantity.
- Applying a
capacityMultiplierallows translating between pool capacity units and standard resource units, converting one pool core into two standard CPUs. - The final calculated amount is recorded in the pod status.
# ResourceSlice
apiVersion: resource.k8s.io/v1
kind: ResourceSlice
metadata:
name: native-resource-slice
spec:
driver: dra.example.com
nodeName: my-node
pool: { name: "node-pool", generation: 1, resourceSliceCount: 1 }
devices:
- name: socket0
attributes:
"dra.example.com/type": "socket"
allowMultipleAllocations: true
capacity:
"dra.example.com/cores": "64"
"dra.example.com/memory": "256Gi"
nodeAllocatableResources:
cpu:
mapping:
capacityKey: "dra.example.com/cores"
capacityMultiplier: "2"
memory:
mapping:
capacityKey: "dra.example.com/memory"
capacityMultiplier: "1"
---
# ResourceClaim
apiVersion: resource.k8s.io/v1
kind: ResourceClaim
metadata:
name: shared-cpu-pool-claim
spec:
devices:
requests:
- name: cpu-pool-request
exactly:
deviceClassName: additional-cpu-memory
capacity:
requests:
"dra.example.com/cores": "2"
---
# Pod
apiVersion: v1
kind: Pod
metadata:
name: hpc-workload-pod
spec:
containers:
- name: app
resources:
requests:
cpu: "1"
claims:
- name: cpu-claim
resourceClaims:
- name: cpu-claim
resourceClaimName: shared-cpu-pool-claim
status:
additionalNodeAllocatableResources:
- source:
apiGroup: resource.k8s.io
kind: ResourceClaim
name: shared-cpu-pool-claim
containers:
- app
mapping:
- name: cpu
quantity: "4" # Derived from consumed pool cores (2 cores * multiplier 2)
- Accelerator with Node Allocatable Resource Overhead Shared Across Multiple Containers
- The device publishes auxiliary resource overheads incurred per pod or container reference.
- Specifying both a fixed cost per pod and a variable cost per container allows modeling complex host memory dependencies.
- The scheduler compiles the active referencing containers array to compute the total overhead.
- These overheads accumulate without requiring it be specified inside the pod specification.
# ResourceSlice
apiVersion: resource.k8s.io/v1
kind: ResourceSlice
metadata:
name: my-node-xpus
spec:
driver: xpu.example.com
nodeName: my-node
devices:
- name: xpu-model-x-001
attributes:
example.com/model: "model-x"
nodeAllocatableResources:
memory:
overhead:
perPod: "1Gi"
perContainer: "500Mi"
---
# ResourceClaim
apiVersion: resource.k8s.io/v1
kind: ResourceClaim
metadata:
name: tensor-accelerator-claim
spec:
devices:
requests:
- name: xpu-request
exactly:
deviceClassName: ai-accelerators
count: 1
---
# Pod
apiVersion: v1
kind: Pod
metadata:
name: ml-inference-pod
spec:
containers:
- name: app-c1
resources:
claims:
- name: gpu-ref
- name: app-c2
resources:
claims:
- name: gpu-ref
resourceClaims:
- name: gpu-ref
resourceClaimName: tensor-accelerator-claim
status:
additionalNodeAllocatableResources:
- source:
apiGroup: resource.k8s.io
kind: ResourceClaim
name: tensor-accelerator-claim
containers:
- app-c1
- app-c2
overhead:
- name: memory
perPod: "1Gi"
perContainer: "500Mi"
- Partitionable Devices
- The resource is modeled hierarchically across NUMA or cache boundaries using shared counter sets.
- The specific capacity consumed from the shared counter set determines the direct resource footprint.
# ResourceSlice
apiVersion: resource.k8s.io/v1
kind: ResourceSlice
metadata:
name: cpu-topology-slice
spec:
driver: dra.example.com
nodeName: my-node
sharedCounters:
- name: node-cpu-counters
counters:
"dra.example.com/cpu": { value: "32" }
devices:
# NUMA Level Devices
- name: numa-0
attributes:
dra.example.com/type: numa
dra.example.com/numaID: "0"
capacity:
"dra.example.com/cpu": "16"
consumesCounters:
- counterSet: node-cpu-counters
counters:
"dra.example.com/cpu": "16"
nodeAllocatableResources:
cpu:
mapping:
capacityKey: "dra.example.com/cpu"
capacityMultiplier: "1"
# L3 Cache Level Devices
- name: numa-0-l3-0
attributes:
dra.example.com/type: l3cache
dra.example.com/numaID: "0"
dra.example.com/l3ID: "0"
capacity:
"dra.example.com/cpu": "8" # L3 cache drawing 8 CPUs
consumesCounters:
- counterSet: node-cpu-counters
counters:
"dra.example.com/cpu": "8"
nodeAllocatableResources:
cpu:
mapping:
capacityKey: "dra.example.com/cpu"
capacityMultiplier: "1"
- name: numa-0-l3-1
attributes:
dra.example.com/type: l3cache
dra.example.com/numaID: "0"
dra.example.com/l3ID: "1"
capacity:
"dra.example.com/cpu": "8"
consumesCounters:
- counterSet: node-cpu-counters
counters:
"dra.example.com/cpu": "8"
nodeAllocatableResources:
cpu:
mapping:
capacityKey: "dra.example.com/cpu"
capacityMultiplier: "1"
# ... additional devices for numa-1
---
# ResourceClaim
apiVersion: resource.k8s.io/v1
kind: ResourceClaim
metadata:
name: l3-cache-claim
spec:
devices:
requests:
- name: l3-req
exactly:
deviceClassName: dra-l3-caches
count: 1
---
# Pod
apiVersion: v1
kind: Pod
metadata:
name: pod1
spec:
containers:
- name: fast-app
resources:
claims:
- name: cache-claim
resourceClaims:
- name: cache-claim
resourceClaimName: l3-cache-claim
status:
additionalNodeAllocatableResources:
- source:
apiGroup: resource.k8s.io
kind: ResourceClaim
name: l3-cache-claim
containers:
- fast-app
mapping:
- name: cpu
quantity: "8" # Derived from specific consumed capacity key of the L3 cache device
- Fungible Resource Claim (GPU or CPU)
- The claim template uses
firstAvailableto request either a GPU or a slice of 30 exclusive CPUs. - If the scheduler selects the GPU,
additionalNodeAllocatableResourcesremains empty because the GPU does not manage node allocatable resources. - If the scheduler selects the CPU slice,
additionalNodeAllocatableResourcesis populated with the 30 CPUs.
# ResourceClaimTemplate for Fungibility
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
name: gpu-or-cpu-template
spec:
spec:
devices:
requests:
- name: gpu-or-cpu-req
firstAvailable:
- name: gpu
deviceClassName: gpu-class
count: 1
- name: cpu
deviceClassName: cpu-class
capacity:
requests:
"dra.example.com/cpu": "30"
---
# Pod
apiVersion: v1
kind: Pod
metadata:
name: fungible-pod
spec:
containers:
- name: my-app
resources:
requests: { cpu: "1", memory: "1Gi" }
claims: [{ name: "gpu-or-cpu" }]
resourceClaims:
- name: gpu-or-cpu
resourceClaimTemplateName: gpu-or-cpu-template
---
# Pod Status (Scenario A: GPU Selected)
status:
additionalNodeAllocatableResources: []
---
# Pod Status (Scenario B: CPU Selected)
status:
additionalNodeAllocatableResources:
- source:
apiGroup: resource.k8s.io
kind: ResourceClaim
name: fungible-pod-gpu-or-cpu-xyz12
containers: ["my-app"]
mapping:
- name: cpu
quantity: "30"
API Validation
- The keys in the
nodeAllocatableResourcesmap must be exactlycpu,memory,hugepages-<size>, orephemeral-storage. All other names, including extended resources, are rejected. - Within a single resource mapping, at least one of the
mappingoroverheadfields must be specified. - If
mappingis specified, it must use eitherdeviceMultiplieror a combination ofcapacityKeyandcapacityMultiplier. These options are mutually exclusive. - If
capacityKeyis specified, it must be a valid qualified name andcapacityMultiplieris required. - If the
overheadfield is specified, it must contain at least one non-negative value for either theperPodorperContaineroverhead quantities. - For
PodStatusupdates, each entry in theadditionalNodeAllocatableResourcesarray is validated as follows:sourcemust be unique across all entries inadditionalNodeAllocatableResources.source.apiGroupandsource.kindmust be known types. Currently, onlyapiGroup: "resource.k8s.io"andkind: "ResourceClaim"is supported (future KEPs may introduce additional source types along with their corresponding validation rules).- When
source.kindisResourceClaim,source.namemust reference a validResourceClaimassociated with the pod (inpod.spec.resourceClaims,pod.status.resourceClaimStatuses, orpod.status.extendedResourceClaimStatus), and every container listed incontainersmust exist inpod.specand reference that claim. mappingandoverheadentries must contain valid node allocatable resource names and non-negative resource quantities.
- Ephemeral storage is validated and supported for root filesystem allocations (see Ephemeral Storage Support and Eviction in Node Resource Enforcement for details).
Kube-Scheduler Changes
The scheduler stages walkthrough below focuses on the DRA node allocatable use case which includes DynamicResources and NodeResourcesFit plugins, but
both the scheduler architecture and the pod.status.additionalNodeAllocatableResources API are designed so that other existing/future plugins
can contribute additional node allocatable footprints on a pod. To support this, the scheduler design follows the below principles:
- Per-node
AdditionalNodeAllocatableResourcesare stored inCycleStatewith a shared, plugin-agnostic key (outside any individual plugin’s state). This enables multiple plugins to incrementally append or read a pod’s additional node allocatable resources for each candidate node. - Each plugin that contributes or checks node allocatable resources in the
Filterstage must evaluate the entire resource footprint computed so far (by reading other plugins contributions fromCycleState). This ensures that the last plugin to run in the chain always validates the complete footprint regardless of plugin execution order. NodeResourcesFitplugin will remain the single owner for scoring and would also be the single writer that patches the pod status (pod.status.additionalNodeAllocatableResources) duringPreBind.
For the DRA node allocatable case, the scheduling stages work as follows:
PreFilter Stage: Neither plugin changes its
PreFilterbehavior- NodeResourcesFit Plugin: Continues to calculate and cache only the pod’s standard
pod.specresource requests inCycleStatewithout filtering nodes. - DynamicResources Plugin: Continues to validate the
ResourceClaimand its associatedDeviceClass. It ensures that the referenced classes exist.
- NodeResourcesFit Plugin: Continues to calculate and cache only the pod’s standard
Filter Stage: This stage performs the node level checks to determine if a pod fits on a specific node.
- NodeResourcesFit Plugin: This plugin checks the pod’s demand against remaining node capacity:
- If
NodeResourcesFitruns before theDynamicResourcesplugin (the default order), the DRA values are not resolved yet, so it checks resource fit based on spec requests only. TheDynamicResourcesplugin then does the authoritative check based on spec + DRA later in the chain. - If
NodeResourcesFitruns after theDynamicResourcesplugin, the pod’s node specific DRA allocations are already resolved inCycleState, so it includes them and checks fit based on spec + DRA. In both orders, the last check to run covers the full demand, so the resource fit result does not depend on the plugin order.
- If
- DynamicResources Plugin: This plugin performs the combined check (spec + DRA resources) only if any of the pod’s
ResourceClaims request node allocatable resources.- The plugin tries to allocate devices to all the resource claims of the pod.
- Claim Resource Calculation: For each allocated device, the plugin reads
spec.devices[*].nodeAllocatableResourcesfrom the node’sResourceSliceand computes the quantity for each resource based on whethermappingand/oroverheadis specified:- If
mappingis specified, the quantity is derived using thecapacityKey,capacityMultiplier, ordeviceMultiplierfields.- If
capacityKeyis set, the base quantity is the consumed capacity from the claim allocation results multiplied bycapacityMultiplier. - If
capacityKeyis omitted, thedeviceMultiplieris applied directly to the count of allocated devices.
- If
- If
overheadis specified, the auxiliary overhead is calculated by summing anyperPodcost and the variableperContainercost scaled by the number of active container references.
- If
- The plugin calculates the total effective demand for each node allocatable resource by:
- Summing up container requests from the pod spec requests and the amounts determined from DRA claims.
- If a claim is referenced by multiple containers, the resource values are accounted for only once.
- If pod level resources are also specified, that takes precedence and determines the resource footprint of the pod.
- With DRA prioritized lists, specifying the same resource at the pod level (
pod.spec.resources) fixes the pod’s footprint regardless of which subrequest is selected later by the DRA plugin, so container-level requests (without pod-level resources) are recommended when the DRA footprint varies across prioritized list options.
- With DRA prioritized lists, specifying the same resource at the pod level (
- Validation: The plugin validates the following scenarios:
- If pod level resources are specified, the plugin will validate that the sum of effective
requests (standard + DRA claims) does not exceed the budget set at the pod level in
pod.spec.resources(details). - The plugin enforces sharing rules based on mapping. If a claim is already assigned to
an existing pod and the allocated device uses direct device mappings
(
nodeAllocatableResources[...].mapping), shared access is blocked across pods to prevent cgroup conflicts. Auxiliary overhead mappings (nodeAllocatableResources[...].overhead) are allowed to share across pods (details).
- If pod level resources are specified, the plugin will validate that the sum of effective
requests (standard + DRA claims) does not exceed the budget set at the pod level in
- This total effective demand is checked against the node’s allocatable resources and node is filtered out if it does not have enough capacity.
- The calculated node allocatable resource allocations for the pod on this specific node (
AdditionalNodeAllocatableResources) are stored inCycleStatewith a plugin-agnostic key. This is needed for passing the node-specific allocation information to other plugins’Filterstage or to the laterScore,Assume, andPreBindstages.
- NodeResourcesFit Plugin: This plugin checks the pod’s demand against remaining node capacity:
Score Stage:
NodeResourcesFitandNodeResourcesBalancedAllocationinclude the pod’s node specific DRA allocations when scoring a candidate node. This is done by reading theAdditionalNodeAllocatableResourcesstored in the sharedCycleStatekey duringFilter(see Scoring).Scheduler Internal Cache Update: After a node is selected, the scheduler updates its internal cache to reflect the resources consumed by the new pod. This stage is critical for maintaining the internal cache consistent. The scheduler framework “assumes” the pod will run on the selected node and updates its cache without waiting for bind (updating the API server) to succeed. Without an “assume” step, the scheduler might try to place other pods on the same node using stale resource information, potentially leading to oversubscription. The Assume phase reserves the resources in the scheduler’s in-memory cache immediately.
- The scheduler framework retrieves the node-specific allocation status from the cycle state which was populated
during the
Filterstage. - This is then applied to the in-memory copy of the Pod object’s status (
pod.status.additionalNodeAllocatableResources) that the scheduler is about to “assume”. - The pod’s overall resource footprint is natively computed via
PodInfo.CalculateResource()(pkg/scheduler/framework/types.go), which sums standard requests and allocations frompod.status.additionalNodeAllocatableResources. This is added tonodeInfo.Requested.
- The scheduler framework retrieves the node-specific allocation status from the cycle state which was populated
during the
PreBind Stage: This stage performs actions right before the pod is bound to the node.
NodeResourcesFitimplementsPreBind(andUnreserve) to read the selected node’s resolved node allocatable allocations fromCycleStateand patchpod.status.additionalNodeAllocatableResourceswhich is evaluated byResourceQuotaadmission inkube-apiserver—see Enforcement for Incoming Pods with DRA Claims and consumed by Kubelet during pod admission and cgroup enforcement. TheDynamicResourcescallsbindClaimto updateResourceClaim.Statuswith the allocated devices and reservation.- If
NodeResourcesFitruns before theDynamicResourcesplugin (the default order),NodeResourcesFit.PreBindpatchespod.status.additionalNodeAllocatableResourcesfirst. If quota is exceeded,PreBindfails andResourceClaimis not reserved. IfNodeResourcesFit.PreBindsucceeds andDynamicResources.PreBind(orBind) fails subsequently,NodeResourcesFit.Unreserveclearspod.status.additionalNodeAllocatableResourcesto release the quota. - If
NodeResourcesFitruns after theDynamicResourcesplugin,DynamicResources.PreBind(bindClaim) runs first, followed byNodeResourcesFit.PreBindpatchingpod.status.additionalNodeAllocatableResources. IfNodeResourcesFit.PreBindfails (e.g., on quota rejection), the scheduler runsUnreserve, whereDynamicResources.Unreserveremoves the claim reservation (ReservedFor) and reverts the claim allocation.
- If
Bind Stage: This stage executes asynchronously after the main scheduling cycle has decided on a node. The scheduler listens for pod
Updateevents, and transitions the pod from the “assumed” state to “bound” if the bind process succeeded. The resource accounting on theNodeInfodoes not change at this point (as they were previously accounted for during the “Assume” step). If the bind fails, or if the Kubelet later rejects the Pod, the scheduler detects this and reverts the resource allocation in its cache, decrementingnodeInfo.Requested.
Requeueing: When a pod cannot be scheduled, the scheduler only checks QueueingHints for the specific plugin that rejected the pod in Filter.
Because either NodeResourcesFit or DynamicResources (or a future plugin) could be the one that rejects a pod for insufficient node capacity,
each of them must listen for events that free up node capacity
- NodeResourcesFit Plugin: Already registers pod-deletion, pod-scale-down, and node-capacity update events.
- DynamicResources Plugin: This will be updated to register pod-deletion and pod-scale-down events (in addition to its existing claim, slice, class, and node events) so pods rejected for DRA included capacity are requeued when capacity frees up instead of waiting for the periodic unschedulable-queue flush.
Resource Calculation
To ensure consistent resource accounting across multiple consumers, the core logic for calculating a pod’s total
resource footprint, including DRA-managed node allocatable resources, will be centralized in the PodRequests function within the
k8s.io/component-helpers/resource package. This helper function is currently used by various components, including scheduler plugins like NodeResourcesFit, the NodeInfo cache update, and the Kubelet’s admission handler.
Reader components (such as HPA, ResourceQuota, and kubectl describe node) can consume pod.status.additionalNodeAllocatableResources via PodRequests()
without checking a feature gate, whereas components that run or simulate scheduling (kube-scheduler, Cluster Autoscaler) will still check the DRANodeAllocatableResources feature gate.
The total node allocatable resource requirements for a pod are determined as follows:
- With Pod-Level Resources: If pod-level resources (
pod.spec.resources.requests) are specified for a resource, they define the overall footprint for that resource. Individual container-level requests and any DRA status allocations/overheads are ignored. - Without Pod-Level Resources: The footprint is calculated by combining standard container requests and DRA status allocations:
- For each container, its effective request is the sum of its standard resource requests and any DRA allocations it references. We get these
DRA allocations from the fields in
pod.status.additionalNodeAllocatableResources(bothmappingandoverheadmappings). - If init containers reference a claim with an overhead.perContainer mapping, we rely on the existing logic used with standard requests where the peak of regular and init containers’ resources is considered.
- Any pod-scoped DRA overheads (
overhead.perPod) are added directly to this total.
- For each container, its effective request is the sum of its standard resource requests and any DRA allocations it references. We get these
DRA allocations from the fields in
- Pod Overhead: In both cases, if standard pod overhead (
pod.spec.overhead) is specified, it is added to the final calculated sum. - Interaction with In-Place Resizing:
- With Pod-Level Resources:
- When a running pod is resized, the pod-level Spec (
pod.spec.resources.requests) is updated. Before the Kubelet accepts and actuates this resize, the scheduler computes the footprint (inPodRequests()) using the maximum ofdesired(pod.spec.resources.requests),allocated(pod.status.allocatedResources), andactuated(pod.status.resources.requests) resources. - Because the pod-level
allocatedandactuatedstatus APIs are updated to include DRA, thismaxcalculation automatically accounts for the DRA resources. We do not need to includepod.status.additionalNodeAllocatableResourcesagain.
- When a running pod is resized, the pod-level Spec (
- Without Pod-Level Resources:
- When a running pod is resized, standard container requests are updated in the Spec. Before Kubelet actuates the resize,
PodRequests()computes the standard container requests using the maximum ofdesired(container.resources.requests),allocated(containerStatuses[*].allocatedResources), andactuated(containerStatuses[*].resources.requests) resources, and adds the static DRA resources. Sinceactuatedalready contains DRA enforced values, we need to deduplicate this before addingpod.status.additionalNodeAllocatableResourcesso that DRA resources are not double-counted.
- When a running pod is resized, standard container requests are updated in the Spec. Before Kubelet actuates the resize,
- With Pod-Level Resources:
Integration with Pod Level Resources
When Pod Level Resources are specified (pod.spec.resources), it continues to set the overall budget for the pod.
Node allocatable resources added to individual containers via DRA claims must be accounted for within this pod-level budget.
The effective resource request for a container is the sum of its base request specified in spec.containers[].resources.requests
and any additional resources allocated through DRA claims.
Currently, with pod level resources, an admission time validation ensures that the sum of container requests does not
exceed pod level requests. However, this is insufficient for pods with node allocatable resource claims, as their exact quantities
are only determined after the DynamicResources scheduler plugin allocates devices. This allocation can be dynamic,
especially for claims with prioritized lists (fungibility use cases).
Therefore, the DynamicResources plugin must perform an additional validation step during its Filter stage. After allocating
devices to claims and calculating the node allocatable resources added, the plugin will verify that the total effective pod demand
(standard container requests + DRA node allocatable resources) does not surpass the limits set in pod.spec.resources.
If a pod requests a specific set of devices via DRA claims, and the resulting node allocatable resource footprint
(base container + DRA additions) exceeds the pod.spec.resources budget, this failure is global to the pod.
The DynamicResources plugin would return UnschedulableAndUnresolvable.
Note: DRA Prioritized Lists (Fungibility) Limitation: Because pod level resources acts as a strict ceiling, using prioritized lists with pod level resources is a known limitation. The pod level budget must be sized to fit the maximum resource option in the prioritized list. If the scheduler chooses a lower-overhead option, the capacity remains unused. It is not recommended to use prioritized lists with pod level resources.
Handling Shared Claims
Intra-Pod Sharing:
Containers within the same pod can reference the same ResourceClaim. The node allocatable resources associated with the claim are accounted for
only once for the entire pod, as described in the Resource Calculation section. The resource calculation shared library function
PodRequests() can effectively handle de-duplication for claims shared within a single pod, as all necessary information is self-contained
within the Pod scope (standard requests in Spec and DRA requests in status.additionalNodeAllocatableResources).
Inter-Pod Sharing:
Sharing ResourceClaims that manage node allocatable resources across different pods is evaluated differentially depending on the mapping type established in the Device mapping:
- CPU/Memory Direct Mappings (
Mappingfield is set): TheDynamicResourcesplugin continues to block sharing across pods (returningUnschedulableAndUnresolvable). Sharing pools of direct native resources creates severe accounting ambiguities (attributing fractional pool costs against distinct pod-level budgets) and intense Kubelet cgroup reconciliation friction. - Accelerator Overheads (Only
Overheadfield is set): TheDynamicResourcesplugin allows sharing across pods. Auxiliary overheads represent host memory or auxiliary tracking structures required per consumer pod/reference. Because these represent standard additive overheads without dynamic draw-down interactions, the scheduler and Kubelet safely accumulate and sum all resources directly frompod.Status.AdditionalNodeAllocatableResourcesfor each individual pod independently.
The DynamicResources plugin enforces the sharing restriction during the Filter stage by inspecting the claim’s
existing consumers (claim.status.reservedFor): if an allocated claim with a mapping entry is already reserved for
another pod, the node is rejected with UnschedulableAndUnresolvable. A mapped claim has exactly one consuming pod,
and this also holds within a pod group: a claim reserved for a whole group (e.g., gang scheduling) still cannot be
shared by its members, because each member pod would record the full mapped quantity in its own status and one
physical device would be counted once per member. Overhead-only claims remain shareable, inside and outside pod
groups.
No new scheduler framework API is required for this. The NodeAllocatableDRAClaimState type and the corresponding
NodeInfo tracking introduced in the initial alpha (v1.36) were removed in the alpha2 rework of the
k8s.io/kube-scheduler staging module.
DRA Admin Access
Admin access (adminAccess: true on a claim request) is a privileged mode for monitoring and diagnostics, only allowed in
namespaces labeled resource.kubernetes.io/admin-access: "true". In the scheduler, an admin access allocation does not consume
the device: the allocator can hand out a device that is already allocated to a workload, and the admin allocation does not block
later allocations. Each such device is marked with adminAccess: true in the allocation result.
For node allocatable resources, the workload’s own claim already accounts for the device’s resources, so charging the admin
claim again would double count the node. Admin access results therefore contribute no footprint, get no entry in
pod.status.additionalNodeAllocatableResources, and do not block or get blocked by the mapped-claim sharing rule.
Multiple Claims per Container
A single container can reference multiple DRA claims. The node allocatable resources from each distinct claim are summed up to contribute to the pod’s total resource requirements.
Example:
- Combining additive policies.
ClaimA - requests 4 CPUs
ClaimB - requests 2 CPUs
- Pod 1
- Container “c1”
- Spec: requests 1 CPU
- claims: ClaimA, ClaimB
- Container “c2”
- Spec: requests 2 CPU
- claims: ClaimA
- Result:
- Pod Effective CPU = 1 (c1 PodSpec) + 4 (ClaimA) + 2 (ClaimB) + 2 (c2 PodSpec) = 9 CPUs.
- Claim A is accounted for only once
- Pod 1
Unreferenced Claims
If a ResourceClaim is listed in pod.spec.resourceClaims but not referenced by any container in pod.spec.containers[*].resources.claims,
the resources associated with this claim are still accounted for against the node’s capacity once. This is because
the DRA allocator allocates the devices to the claim making them unavailable to others (e.g., exclusive CPUs requested through a claim).
This will be enforced in the PodRequests() helper function when computing the pod resource footprint.
Scoring
Currently, resource based scoring plugins (NodeResourcesFit and NodeResourcesBalancedAllocation) evaluate
candidate nodes by comparing a pod’s resource requests against each node’s allocatable capacity and existing usage
(nodeInfo.GetRequested()). Because pod.Spec requests are uniform across all nodes, the scheduler computes the
pod’s request vector once during PreScore and reuses it to score every node.
With DRA node allocatable resources, scoring accounts for dynamic claim resources for this calculation:
- Existing pods on a node: Covered automatically. Their footprint in
nodeInfo.GetRequested()includes node allocatable resources allocated to claims viapod.status.additionalNodeAllocatableResources(both for running pods and assumed pods in the scheduler cache). - The pod being scored: The footprint can vary per node because the same claim may resolve to different resource
quantities on different nodes. During
Filter, theDynamicResourcesplugin records the pod’s node-specific claim allocations inCycleState. DuringScore, bothNodeResourcesFitandNodeResourcesBalancedAllocationread this allocation fromCycleStatefor the candidate node and include it in the pod’s request vector. Pods with no such claims, and nodes where the plugin did not run, use the requests computed inPreScoreunchanged. No device resolution is added to the scoring path.
Preemption
If a high-priority Pod is unschedulable due to insufficient resources, the scheduler tries to find a suitable node by preempting lower-priority pods:
- The default preemption plugin simulates evicting (
SelectVictimsOnNode()) lower-priority pods. Because the victim pods are already running on the node, and the pod status is populated with DRA allocations, the resource calculation helper function (PodRequests()) accurately subtracts both the victim’s Spec requests and its dynamic status claim allocations. - When the default plugin simulates adding back candidate victims one by one to see if the incoming pod still fits, this check automatically aggregates both standard Spec requests and dynamic status claim allocations for the reprieved pods.
- During these eviction and reprieve simulations, the preemption plugin always checks (RunFilterPluginsWithNominatedPods()) if the pod fits. The dynamic resources plugin node-fit check includes DRA allocations, the preemption plugin correctly identifies candidate nodes.
- There is an independent proposal for DRA preemption. However, because node allocatable claims are mapped to standard resources and are already included in the scheduler resource footprint calculation and internal cache updates, DRA-based node allocatable requests are automatically considered during preemption even without the DRA preemption feature enabled.
Node Resource Enforcement and Isolation
Scope
The Kubelet’s primary responsibility is to set up the cgroup hierarchy, set pod-level ceilings, and container-level headroom (limits). It guarantees that the
pod-level parent cgroup bounds have the correct resource ceilings, and container-level cgroups have safe defaults (e.g., CFS quota, memory limits) so that workloads
can utilize their claim resources without throttling or OOM kills. DRA drivers can then modify these container-specific settings configured by the Kubelet or apply
new enforcements (e.g., CPU pinning or binding memory to specific NUMA nodes) by interfacing directly with the Container Runtime (e.g., a CPU DRA driver using NRI to set cpuset.cpus).
Considering DRA resources in Kubelet cgroup enforcement guarantees that any container-level modifications or overrides applied by a DRA driver are contained and
cannot affect other co-located pods on the node. This helps to keep the KEP generic and independent of specific DRA driver implementations.
Key Principles
- Kubelet’s DRA-specific adjustments to cgroup enforcement are derived solely from
pod.status.additionalNodeAllocatableResourcesas updated by the scheduler. - If Pod Level Resources are explicitly specified, that takes precedence at both the scheduler level for accounting and the node level for cgroup enforcement.
- The QoS classification of a pod remains determined strictly by the standard requests and limits in the PodSpec. DRA claims do not alter the pod’s QoS tier.
- If a standard request or limit is not specified in the spec, the defaulting mechanism that we currently have (for example, setting CPU shares to 2, or quota to unlimited) remains true. The defaulting logic at the pod level and container level cgroups is still determined based on standard Spec, and DRA does not change that.
Cgroup Enforcement
To enforce container and pod-level cgroup settings, Kubelet reads AdditionalNodeAllocatableResources from pod.Status and uses this
information along with standard resource requests and limits specified in the Pod Spec (pod.spec.containers[].resources and pod.spec.resources
when using Pod-Level Resources) to determine the overall cgroup allocations. Kubelet evaluates cgroup settings at both the pod level and container level.
Workload resource boundaries are actuated at two distinct levels in the host cgroup v2 hierarchy:
- Pod-Level parent cgroups
- Establish the overall aggregate resource boundary for the entire pod.
- This parent cgroup acts as a shared pool of resources, enabling containers to dynamically share CPU and memory while safely bounding the pod’s overall resource footprint.
- Enforced directly by kubelet.
- Container-Level cgroups
- Applies granular resource isolation boundaries directly to the container based on container Spec (or default values when not specified).
- Enforced through CRI.
Kubelet translates Pod Spec resource requests and limits into corresponding cgroup settings using these core cgroup properties:
- CPU Requests are mapped to CPU Shares/Weight (
cpu.weight): Controls the relative CPU scheduling weight/priority of the pod or container when the node experiences CPU contention. - CPU Limits are mapped to CPU Quota (
cpu.max): Caps the absolute maximum CPU time the pod/container can consume in a time window (configurable). - Memory Limits are mapped to Memory Limit (
memory.max): Caps the absolute maximum memory (RAM) the pod/container can consume. - HugePages Limits are mapped to HugePages Limit (
hugetlb.<size>.max): Caps the maximum hugepage allocation size.
Kubelet also sets up the cgroup directories for the pod based on the QoS class (Guaranteed, BestEffort or Burstable). DRA based allocation does not
have an influence on the QOS class of the pod and how Kubelet sets up cgroup hierarchies.
Kubelet evaluates cgroup settings at both the pod level and container level as follows:
Pod-Level Cgroup Settings
Without DRA:
If PodLevelResources are enabled and explicitly specified (pod.spec.resources.requests and pod.spec.resources.limits), Kubelet sets the pod-level cgroup settings
exactly to those explicit values. If PodLevelResources are not specified, Kubelet sums up all container-level requests and limits and sets the pod level cgroup settings.
With DRA:
If PodLevelResources are enabled and explicitly specified (pod.spec.resources.requests and pod.spec.resources.limits), Kubelet sets the pod-level cgroup settings exactly
to those explicit values without adding DRA allocations. If PodLevelResources are not specified, Kubelet sums up all container-level requests and limits and adds DRA allocations.
At the pod level, Kubelet sets the cgroup parameters as follows:
CPU Shares = MilliCPUToShares( Sum(Spec.Requests[cpu]) + DRADirectMapped(cpu) + DRAOverheadMappedPodTotal(cpu) )
CPU Quota = Sum(Spec.Limits[cpu]) + DRADirectMapped(cpu) + DRAOverheadMappedPodTotal(cpu)
Memory Limit = Sum(Spec.Limits[memory]) + DRADirectMapped(memory) + DRAOverheadMappedPodTotal(memory)
HugePages Limit = Sum(Spec.Limits[hugepages-<size>]) + DRADirectMapped(hugepages-<size>) + DRAOverheadMappedPodTotal(hugepages-<size>)
Sum(Spec.Requests[resource]): Sum of requests across all containers in the pod.Sum(Spec.Limits[resource]): Sum of limits across all containers in the pod.DRADirectMapped(resource): Sum of direct mapped DRA allocations for all the claims referenced in the pod (obtained frompod.status.additionalNodeAllocatableResources[].mapping[].quantity).DRAOverheadMappedPodTotal(resource): Sum of overhead mapped DRA allocations across all distinct claims allocated to the pod, obtained asPerPod + (PerContainer * len(containers)).
Why Pod Level Cgroup Limits includes DRA allocations?
- The pod’s cgroup slice establishes the absolute upper ceiling (
cpu.max,memory.max,hugetlb.<size>.max) for the entire pod workloads footprint. - If DRA allocations (direct or overhead) are not added to the pod workloads cgroup limits, the pod-level ceiling remains locked at standard Spec-pure limits The moment any container attempts to utilize its DRA capacity, the overall pod usage will hit the uninflated parent boundary, resulting in immediate CPU throttling, memory OOM kills, or hugepage allocation failures.
- If
PodLevelResourcesare explicitly declared inpod.spec.resources.limits, the Kubelet respects the user’s aggregate pod limits budget and does not add DRA allocations, expecting the user to have configured the pod level settings to include DRA allocations.
Why Pod Level Requests / CPU Shares includes DRA allocation ?
- Since DRA CPU resources are accounted during node capacity calculations during scheduling, the scheduler has already reserved and deducted these CPUs from the node’s capacity. Including the DRA values at the pod-level cgroup ensures that the host kernel actually honors this scheduler-level resource reservation under node contention.
- In Linux, CPU shares (
cpu.weight) act as relative priority weights that are only enforced when the entire node experiences heavy CPU contention. Including DRA requests at the pod level ensures the entire pod successfully secures its aggregate resource footprint against other pods on the node. Including the DRA values at the pod-level cgroup ensures that the host kernel actually honors this scheduler-level resource reservation under node contention.- Example: If a container requests
100mCPU through a standard request, and gets1 CPUthrough a DRA claim for a GPU device (overhead), setting the CPU shares only based on the standard100mCPU request would starve the container during node CPU contention.
- Example: If a container requests
- This is in line with the scope of the KEP that Kubelet sets the pod-level cgroup boundaries based on DRA and sets safe defaults at the container level allowing for the DRA driver to modify. This allows for the DRA drivers to model both shared and exclusive resources.
Container-Level Cgroup Settings
At the container level, Kubelet sets the cgroup parameters as follows:
CPU Shares = MilliCPUToShares(Spec.Requests[cpu]) # No changes
CPU Quota = Spec.Limits[cpu] + DRADirectMapped(cpu) + DRAOverheadMappedPerContainer(cpu) + DRAOverheadMappedPerPod(cpu)
Memory Limit = Spec.Limits[memory] + DRADirectMapped(memory) + DRAOverheadMappedPerContainer(memory) + DRAOverheadMappedPerPod(memory)
HugePages Limit = Spec.Limits[hugepages-<size>] + DRADirectMapped(hugepages-<size>) + DRAOverheadMappedPerContainer(hugepages-<size>) + DRAOverheadMappedPerPod(hugepages-<size>)
Spec.Requests[resource]: Standard request specified inpod.spec.containers[].resources.requests(or default value if unset)Spec.Limits[resource]: Standard limit specified inpod.spec.containers[].resources.limits. If container-level limits are omitted butPodLevelResources(pod.spec.resources.limits) are explicitly specified, this value falls back to the pod level resource limit.DRADirectMapped(resource): Sum of direct compute resources allocated via DRA (e.g., resources allocated via cpu/memory dra driver), obtained frompod.status.additionalNodeAllocatableResources[].mapping[].quantity.DRAOverheadMappedPerContainer(resource): Sum of overhead resources allocated via DRA (e.g., additional cpu/memory resources for a GPU device), obtained frompod.status.additionalNodeAllocatableResources[].overhead[].perContainer.DRAOverheadMappedPerPod(resource): Sum of overhead DRA allocations for the pod, obtained frompod.status.additionalNodeAllocatableResources[].overhead[].perPod.- Since the claim resources are shared by all containers referencing the claim, the per-pod overhead is included in the limit of all the containers, but is counted exactly once at the parent pod-level cgroup ceiling.
Why Container Level Limits includes DRA allocations?
- To allow containers to successfully consume and utilize their allocated DRA claims, their nested container level cgroup limits must be inflated to accommodate the additional capacity. Without this, the container would be immediately throttled or OOM-killed by its spec-only cgroup boundary, completely rendering the DRA allocations unusable.
Why Container Level Requests / CPU Shares DOES NOT INCLUDE DRA allocations?
The Kubelet lacks the context to know whether a DRA allocation represents exclusive resources or shared capacity. If DRA allocates exclusive CPUs, considering those to determine the shared CPU weight would allow the container to unfairly dominate the shared CPU pool during contention with other containers in the pod that do not use exclusive CPUs.
If a claim is shared by multiple containers within a pod, attempting to split the claim’s request among those referencing containers CPU shares would introduce enforcement complexity and ambiguity. To perfectly set the container level Cgroup settings, we would need to know the exact type of resource allocation made through DRA and can be explored as a future enhancement (Pass Allocation Details from Driver to Kubelet).
The risk here is that the DRA allocations are not added to CPU shares, a container using only a claim and no standard request receives minimal CPU weight (
2), risking starvation during contention within the containers of the pod. However, keeping container-level CPU shares only based on spec is a safe and sufficient default for the alpha implementation due to the following reasons:- Including DRA allocation at the pod level CPU shares provides guarantees and due to the cgroup hierarchy, the pod as a whole gets the shares proportional to scheduler allocated resources.
- This is only relevant if the DRA driver does not allocate exclusive CPUs. If the driver allocates exclusive CPUs, there is no contention with other containers in the pod.
- This risk is fully manageable. The scope is strictly to configure the baseline cgroup settings, which the DRA driver can then modify or optimize.
QoS Class Mismatch Risks
Because a pod’s Quality of Service (QoS) class is determined strictly by the standard container resource definitions in pod.Spec and ignores DRA Status allocations,
workloads can experience degradation because of how cgroups are configured by kubelet. The risks vary based on the pod’s resulting QoS category:
1. Pod Categorized as BestEffort
If the Pod Spec completely omits both requests and limits for both CPU and Memory (either at the pod level in pod.Spec.Resources when using Pod-Level Resources,
or across all containers in pod.Spec.Containers[*]), the pod is classified as a BestEffort QoS class. The risks of a pod with DRA claims being categorized as BestEffort are:
- Kubelet places the pod under
kubepods.slice/kubepods-besteffort.slice/. This parent slice has CPU shares (cpu.weight) set toMinShares (2). Under node-wide CPU contention, the container can be starved because of this parent boundary, regardless of its internal cgroup weight (which can be set by the DRA driver). This CPU starvation risk is only relevant if the workload runs in a shared CPU pool; if the DRA driver allocates exclusive CPU cores and pins the container via cgroup cpuset configurations, CPU shares are completely ignored and the core allocation is fully guaranteed without starvation. - BestEffort pods receive the maximum OOM score adjustment (
1000) and are ranked first for preemption and eviction by the Eviction Manager during memory or disk pressure.
2. Pod Categorized as Burstable
If the Pod Spec specifies any standard CPU or Memory request or limit (either at the pod level in pod.Spec.Resources, or for at least one container in pod.Spec.Containers[*]),
but the pod does not meet the strict requirements for the Guaranteed QoS (i.e., where requests must match limits exactly for both CPU and Memory), the pod is classified as Burstable QoS class.
- Since CPU shares (
cpu.weight) remain based strictly on standard requests, a container requesting a small standard amount but receiving a large allocation via DRA would still have lower CPU shares. Similar to BestEffort, this is not relevant if the DRA driver allocates exclusive CPUs and manages core pinning directly.
Potential Mitigations
Any container-level risks due to Kubelet setting defaults/baseline values not considering exact intent of the claim can be solved at the DRA driver level by updating these base values set by kubelet. However, because the driver is strictly confined to operate at the container level, it cannot modify the parent-level Pod cgroup boundaries.
- Ensure that pods using DRA for CPUs are not classified as BestEffort by specifying a non-zero standard CPU or memory request on one of the containers in
pod.Spec. This promotes the pod to the Burstable QoS tier, moving it out of the BestEffort slice where cgroup values are locked at the parent level. - Use Pod-Level Resources to declare the total aggregate requests (including DRA allocations) at the pod level in
pod.Spec.Resources. This works well only when the claim resources are completely deterministic, and it is not suitable for advanced use cases where the mapping between CPU/Memory and the DRA allocation is not 1:1 (such as modeling L3 caches instead of CPUs directly) or when using a DRA prioritized list where the actual allocation quantity is not known until scheduling time.
Long-Term Mitigation - Explicit QoS Class
A robust long-term solution would be to allow workloads to declare an explicit QoS class directly in the Pod Spec, rather than relying on implicit derivations inside Kubelet. This was also explored as part of KEP-1287 to loosen QoS restrictions during in-place pod resizing. With multiple independent variables now affecting a pod’s resource footprint (standard container specs, Pod-Level Resources, in-place resizing, and now DRA), attempting to implicitly derive the QoS class by coordinating all these inputs is highly complicated and remains a maintenance challenge and exploring explicit QoS class configuration is a more desirable path.
Handling Pod Level Resources
When PodLevelResources is used, the Kubelet’s cgroup enforcement must reconcile explicit pod-level limits with DRA allocations. This requires two specific adjustments:
- Pod-Level Cgroup Ceilings:
If explicit pod-level limits are specified, they determine the overall pod budget. The Kubelet sets the pod’s cgroup ceiling exactly to the specified
pod.spec.resources.limits. It does not add the DRA allocations to the pod-level limit, because the DRA resources are already encompassed within this overall budget. - Container-Level Fallbacks: If a container lacks its own limit, the pod-level limit is applied to the container’s cgroup maximum value.
Handling Missing Limits
When a container omits limits for CPU, Memory, or HugePages, the Kubelet sets cgroup default values or sets it based on pod-level settings:
- CPU and Memory:
- Kubelet defaults the container limit to unlimited.
- Kubelet ignores the DRA allocation values for setting limits.
- HugePages:
- Kubelet defaults the container limit to 0.
- Following the same model as CPU and Memory (default to “unlimited”) for HugePages breaks because by setting a hard limit of zero, we block the container from consuming any hugepages allocated by the DRA driver. If DRA requests HugePages, Kubelet sets the limit to DRA.
Container-Level Cgroup Defaults:
CPU Quota = -1 (unlimited)
Memory Limit = unset (unlimited)
HugePages Limit = DRAMapped(hugepages-<size>) + DRAOverhead(hugepages-<size>)
Pod-Level Cgroup Defaults:
If PodLevelResources are explicitly specified (pod.spec.resources.limits), the pod-level cgroup enforces those absolute limits. If PodLevelResources are not specified, the pod-level cgroup limits inherit the unbounded container defaults, summing up HugePages while deduplicating shared claims:
CPU Quota = -1 (unlimited)
Memory Limit = unset (unlimited)
HugePages Limit = DRAMappedUnique(hugepages-<size>) + DRAOverheadUnique(hugepages-<size>)
Handling Kubelet Disabling Quota with Exclusive CPUs
When a container is allocated exclusive CPUs by Kubelet (using static CPU policy for a Guaranteed QoS pod with integer CPU requests), Kubelet disables
CPU quota enforcement (cpu.max = -1) at both the container and pod levels. This is to prevent unexpected throttling (details in Issue 70585).
With this KEP, this behavior remains the same, with the key distinction that Kubelet natively only checks for exclusive CPUs allocated through its standard static CPU policy.
In the case where exclusive CPU allocation is not managed by Kubelet (i.e., static CPU policy is disabled) but is instead handled independently by a DRA driver, Kubelet lacks visibility into this allocation. Consequently, Kubelet will enforce CFS CPU quotas at both the container and pod levels (if all other conditions for setting quota are met — i.e., all containers have limits set or limits are defined at the pod level).
Risk: With Kubelet enforcing quotas while the DRA driver allocates exclusive physical CPUs, the workload could experience the same throttling issues as in issue 70585.
Current Mitigation: While the DRA driver can use container-level hooks to override Kubelet’s defaults and set the container cgroup to unlimited, it cannot modify Kubelet-managed
pod-level parent cgroups. To mitigate this, the container requesting exclusive CPUs through the DRA claim can skip setting limits in the container spec. Under this configuration, Kubelet’s cgroup manager natively skips quota configuration at both container and pod levels and they remain unlimited (cpu.max = -1).
Potential Long-term Mitigation: A proper long-term solution would involve a better coordination mechanism between Kubelet and the DRA driver to delegate cgroup enforcement
responsibilities and avoid having multiple components configuring the same cgroup settings. It needs more design work to establish this handshake mechanism and is currently out of
scope for the alpha stage of this KEP.
Enforcement Use Case Walkthroughs
- Claim + Standard Request
A pod references a shared CPU claim alongside a standard container request and limit.
# Pod Spec
spec:
containers:
- name: c1
resources:
requests: { cpu: "2", memory: "2Gi" }
limits: { cpu: "4", memory: "4Gi" }
claims: [{ name: "cpu-claim" }]
resourceClaims:
- name: cpu-claim
resourceClaimName: shared-cpu-claim
# Pod Status
status:
additionalNodeAllocatableResources:
- source:
apiGroup: resource.k8s.io
kind: ResourceClaim
name: shared-cpu-claim
containers: ["c1"]
mapping:
- name: cpu
quantity: "5"
- name: memory
quantity: "5Gi"
- Pod Level Cgroup:
cpu.weight(CPU Shares): Set based on standard request + DRA mapping: 2 + 5 = 7 CPUs.cpu.max(CPU Quota): Set based on standard limit + DRA (4 + 5) - 9 CPUs.memory.max(Memory Limit): Set based on standard limit + DRA (4 + 5) - 9 GiB.
- Container Level Cgroup:
- C1
cpu.weight(CPU Shares): Set based on standard request - 2 CPUs.cpu.max(CPU Quota): Set based on standard limit + DRA (4 + 5) - 9 CPUs.memory.max(Memory Limit): Set based on standard limit + DRA (4 + 5) - 9 GiB.
- C1
- Outcome: The container can burst up to 9 CPUs and 9 GiB memory. If the DRA driver allocates exclusive CPUs, the container has sole access to them. The standard request from the container spec comes from the shared pool, by setting shares based on Spec request of 2 ensures inter-pod fairness during contention.
- Only Claim, No Standard Request and Limit Specified
A pod references a CPU claim but specifies no standard requests or limits in its Spec.
# Pod Spec
spec:
containers:
- name: c1
resources:
claims: [{ name: "cpu-claim" }]
resourceClaims:
- name: cpu-claim
resourceClaimName: shared-cpu-claim
# Pod Status
status:
additionalNodeAllocatableResources:
- source:
apiGroup: resource.k8s.io
kind: ResourceClaim
name: shared-cpu-claim
containers: ["c1"]
mapping:
- name: cpu
quantity: "5"
- Pod Level Cgroup:
cpu.weight(CPU Shares): Set based on standard request + DRA mapping: 0 + 5 = 5 CPUs.cpu.max(CPU Quota): -1 (Unlimited).- Container Level Cgroup:
- C1
cpu.weight(CPU Shares): Defaults to default minimum value (2 shares).cpu.max(CPU Quota): -1 (Unlimited).
- C1
- Outcome: The container CPU limit remains unlimited as the values are not set in the spec.
- Multiple Containers Sharing a Claim + Standard Request
Two containers in the same pod share a CPU claim and declare individual standard requests and limits.
# Pod Spec
spec:
containers:
- name: c1
resources:
requests: { cpu: "2", memory: "2Gi" }
limits: { cpu: "4", memory: "4Gi" }
claims: [{ name: "shared-claim" }]
- name: c2
resources:
requests: { cpu: "4", memory: "4Gi" }
limits: { cpu: "8", memory: "8Gi" }
claims: [{ name: "shared-claim" }]
resourceClaims:
- name: shared-claim
resourceClaimName: shared-cpu-claim
# Pod Status
status:
additionalNodeAllocatableResources:
- source:
apiGroup: resource.k8s.io
kind: ResourceClaim
name: shared-cpu-claim
containers: ["c1", "c2"]
mapping:
- name: cpu
quantity: "5"
- name: memory
quantity: "5Gi"
- Pod Level Cgroup:
cpu.weight(CPU Shares): Set based on standard requests sum + DRA mapping: (2 + 4) + 5 = 11 CPUs.cpu.max(CPU Quota): Set based on standard limit sum + DRA counted once (4 + 8 + 5) - 17 CPUs.memory.max(Memory Limit): Set based on standard limit sum + DRA counted once (4 + 8 + 5) - 17 GiB.
- Container Level C1 Cgroup:
- C1
cpu.weight(CPU Shares): Set based on standard request - 2 CPUs.cpu.max(CPU Quota): Set based on standard limit + DRA (4 + 5) - 9 CPUs.memory.max(Memory Limit): Set based on standard limit + DRA (4 + 5) - 9 GiB.
- C2
cpu.weight(CPU Shares): Set based on standard request - 4.cpu.max(CPU Quota): Set based on standard limit + DRA (8 + 5) - 13 CPUs.memory.max(Memory Limit): Set based on standard limit + DRA (8 + 5) - 13 GiB.
- C1
- Outcome: Both containers can burst up to their limit + claim amount individually. Over-subscription of limits is allowed. However, by counting the shared claim only once at the pod-level cgroup ceiling, Kubelet guarantees that if both C1 and C2 burst simultaneously, they cannot collectively exceed the reserved pod-level budget of 17. If the DRA driver allocates exclusive CPUs, both containers have access to all the claim CPUs, but if there is contention, C2 gets higher priority based on shares.
- Pod Level Request and Limit + Shared DRA Claim
A pod defines explicit Pod Level Resources, and two containers share a DRA claim without specifying container-level limits.
# Pod Spec
spec:
resources:
requests: { cpu: "5", memory: "5Gi" }
limits: { cpu: "5", memory: "5Gi" }
containers:
- name: c1
resources:
claims: [{ name: "shared-claim" }]
- name: c2
resources:
claims: [{ name: "shared-claim" }]
resourceClaims:
- name: shared-claim
resourceClaimName: shared-cpu-claim
# Pod Status
status:
additionalNodeAllocatableResources:
- source:
apiGroup: resource.k8s.io
kind: ResourceClaim
name: shared-cpu-claim
containers: ["c1", "c2"]
mapping:
- name: cpu
quantity: "5"
- name: memory
quantity: "5Gi"
- Pod Level Cgroup:
cpu.weight(CPU Shares): Set based on explicit pod request - 5.cpu.max(CPU Quota): Set based on explicit pod limit - 5 CPUs.memory.max(Memory Limit): Set based on explicit pod limit - 5 GiB.
- Container Level Cgroup:
- C1 & C2
cpu.weight(CPU Shares): Defaults to minimal value - 2 CPUs.cpu.max(CPU Quota): Inherited from pod-level limit - 5 CPUs.memory.max(Memory Limit): Inherited from pod-level limit - 5 GiB.
- C1 & C2
- Outcome: Because the containers do not specify their own limits, they inherit the pod-level limit as their container cgroup maximum value. Pod Level Resources act as the absolute maximum overall budget for the pod, DRA allocations must fit within this budget.
- Pod Level Request and Limit + Container Requests and Limits + Shared DRA Claims + Sidecar
A pod defines explicit Pod Level Resources, two regular containers share a DRA claim and define individual limits, and a sidecar runs without container limits.
# Pod Spec
spec:
resources:
requests: { cpu: "8", memory: "8Gi" }
limits: { cpu: "15", memory: "15Gi" }
containers:
- name: c1
resources:
requests: { cpu: "2", memory: "2Gi" }
limits: { cpu: "4", memory: "4Gi" }
claims: [{ name: "shared-claim" }]
- name: c2
resources:
requests: { cpu: "4", memory: "4Gi" }
limits: { cpu: "8", memory: "8Gi" }
claims: [{ name: "shared-claim" }]
initContainers:
- name: sidecar
restartPolicy: Always
# No resources specified for sidecar
resourceClaims:
- name: shared-claim
resourceClaimName: shared-cpu-claim
# Pod Status
status:
additionalNodeAllocatableResources:
- source:
apiGroup: resource.k8s.io
kind: ResourceClaim
name: shared-cpu-claim
containers: ["c1", "c2"]
mapping:
- name: cpu
quantity: "5"
- name: memory
quantity: "5Gi"
- Pod Level Cgroup:
cpu.weight(CPU Shares): Set based on explicit pod request - 8.cpu.max(CPU Quota): Set based on explicit pod limit - 15 CPUs.memory.max(Memory Limit): Set based on explicit pod limit - 15 GiB.
- Container Level Cgroup:
- C1
cpu.weight(CPU Shares): Set based on standard request - 2 CPUs.cpu.max(CPU Quota): Set based on standard limit + DRA (4 + 5) - 9 CPUs.memory.max(Memory Limit): Set based on standard limit + DRA (4 + 5) - 9 GiB.
- C2
cpu.weight(CPU Shares): Set based on standard request - 4.cpu.max(CPU Quota): Set based on standard limit + DRA (8 + 5) - 13 CPUs.memory.max(Memory Limit): Set based on standard limit + DRA (8 + 5) - 13 GiB.
- Sidecar
cpu.weight(CPU Shares): Defaults to minimal value - 2.cpu.max(CPU Quota): Inherited from pod-level limit - 15 CPUs.memory.max(Memory Limit): Inherited from pod-level limit - 15 GiB.
- C1
- Outcome: C1 and C2 calculate their limits by adding the DRA burst to their explicit standard limits (9 and 13 respectively). Because the sidecar omits container limits, it inherits the pod-level limit as its container cgroup maximum value (15). The total aggregate bursting for all containers combined is hard-capped at 15 CPUs.
- Multiple Containers Sharing a Claim with Host Resource Overhead
Two containers in the same pod share a GPU claim that incurs both flat pod-level and variable container-level CPU/Memory overheads.
# Pod Spec
spec:
containers:
- name: c1
resources:
requests: { cpu: "2", memory: "2Gi" }
limits: { cpu: "2", memory: "4Gi" }
claims: [{ name: "shared-gpu" }]
- name: c2
resources:
requests: { cpu: "2", memory: "2Gi" }
limits: { cpu: "4", memory: "8Gi" }
claims: [{ name: "shared-gpu" }]
resourceClaims:
- name: shared-gpu
resourceClaimName: shared-gpu-claim
# Pod Status
status:
additionalNodeAllocatableResources:
- source:
apiGroup: resource.k8s.io
kind: ResourceClaim
name: shared-gpu-claim
containers: ["c1", "c2"]
overhead:
- name: cpu
perPod: "1"
perContainer: "500m"
- name: memory
perPod: "1Gi"
perContainer: "500Mi"
- Pod Level Cgroup:
cpu.weight(CPU Shares): Set based on standard requests sum + DRA overhead: 2(C1 Spec request) + 2(C2 Spec request)+ 1(perPod) + 500m * 2 (perContainer for C1 and C2)- 6 CPUs.cpu.max(CPU Quota): Set based on standard limits sum + DRA overhead: 2(C1 Spec limit) + 4(C2 Spec limit) + 1(perPod) + 500m * 2 (perContainer for C1 and C2): 8 CPUs.memory.max(Memory Limit): Set based on standard limits sum + DRA overhead: 4Gi(C1 Spec limit) + 8Gi(C2 Spec limit) + 1Gi(perPod) + 500Mi * 2 (perContainer for C1 and C2): - 14 GiB.
- Container Level Cgroup:
- C1
cpu.weight(CPU Shares): Set based on standard request - 2 CPUs.cpu.max(CPU Quota): Set based on standard limit + container overhead + pod overhead (2 + 0.5 + 1) - 3.5 CPUs.memory.max(Memory Limit): Set based on standard limit + container overhead + pod overhead (4 + 0.5 + 1) - 5.5 GiB.
- C2
cpu.weight(CPU Shares): Set based on standard request - 2 CPUs.cpu.max(CPU Quota): Set based on standard limit + container overhead + pod overhead (4 + 0.5 + 1) - 5.5 CPUs.memory.max(Memory Limit): Set based on standard limit + container overhead + pod overhead (8 + 0.5 + 1) - 9.5 GiB.
- C1
- Outcome: Both containers can burst up to their individual cgroup quotas (3.5 and 5.5 CPUs respectively) to accommodate container-specific driver overheads and the flat pod overhead when operating alone. However, if both containers execute overhead tasks simultaneously, their combined CPU and memory footprint is hard-capped at the pod-level parent ceilings (8 CPUs and 14 GiB memory).
OOM Score Adjustment with DRA
To manage node stability during Out-Of-Memory (OOM) events, Kubelet applies DRA adjustments while calculating OOM score:
DRA claims are not considered when computing the pod’s QoS class.
Pods classified as
GuaranteedorBestEffortbased on standard Spec continue to receive their static scores (-997and1000), and does not change based on DRA.For pods classified as
Burstable, Kubelet incorporates DRA memory requests to calculate a more protective score.# claimMemory: Total memory quantity allocated to the DRA claim # numContainerReferences: Number of containers in the pod referencing this claim draMemoryShare = claimMemory / numContainerReferences # containerMemReq: Base memory request specified in the container's standard Spec # remainingReqPerContainer: Per-container share of unallocated pod-level resources memory request (0 if PodLevelResources is disabled) effectiveMemReq = containerMemReq + remainingReqPerContainer + draMemoryShare # memoryCapacity: Total physical memory capacity of the host node oomScoreAdjust = 1000 - (1000 * effectiveMemReq / memoryCapacity)If multiple containers share a single DRA memory claim, Kubelet divides the claim’s memory quantity equally among the sharing containers. This equal split is an intentional design simplification as Kubelet cannot dynamically track actual memory distribution between the containers sharing the claim and update the OOM score. This follows the same established pattern with Pod Level Resources (PLR), where pod-level memory requests are distributed equally among containers that omit container-level memory requests.
Kubelet Eviction
Under node pressure, Kubelet ranks eviction candidates by how much a pod’s usage exceeds its requests. Usage already includes
DRA allocations, since they are enforced in the pod’s cgroup. The eviction manager would be updated to compute requests as pod
spec plus DRA allocations (using the shared PodRequests component-helpers function). Without this, a pod whose memory
comes mostly from a claim would be prioritized incorrectly for eviction. Kubelet preemption uses the same requests to decide
how much capacity evicting a pod would free.
Ephemeral Storage Support and Eviction
Scope: The ephemeral-storage key in nodeAllocatableResources covers the same local ephemeral storage a pod can
request through the pod spec (spec.containers[].resources.requests["ephemeral-storage"], see
Local ephemeral storage):
the container writable layers, container logs, and emptyDir volumes backed by the node’s root filesystem.
Unlike cpu, memory, and hugepages, which the kernel contains through cgroups, ephemeral storage has no cgroup controller.
Kubelet enforces it by measurement. The scheduler fits requests against the node’s allocatable ephemeral storage, derived from
the filesystem backing the kubelet root directory. Kubelet periodically measures usage in the container writable layers,
container logs, and local emptyDir volumes, and under node disk pressure ranks eviction candidates by how far their usage
is compared to their requests.
How it works with DRA: A claim grants the pod additional root filesystem storage with the same semantics as
spec.containers[].resources.requests["ephemeral-storage"]:
- The scheduler reserves claim-allocated ephemeral storage against
Node.Status.Allocatable["ephemeral-storage"], preventing root filesystem overcommitment. - Under node disk pressure, eviction ranking includes the DRA amounts in a pod’s requests, so a pod consuming its claim-granted storage is not ranked as exceeding its requests.
- Limit eviction, at both pod level and container level, includes the DRA amounts in the limit it enforces. Like cpu and
memory, the DRA amount is added only when the spec declares an
ephemeral-storagelimit; a pod without a spec limit stays unlimited and is bounded by node-pressure eviction alone. emptyDirsize limit eviction is unchanged. AsizeLimit(spec.volumes[].emptyDir.sizeLimit) bounds that one volume, not the pod. Since a claim raises the pod’s total storage, so a volume that grows more than its ownsizeLimitis still evicted.
The resource quantities are available through pod.status.additionalNodeAllocatableResources like cpu and memory, so ResourceQuota
accounts them under requests.ephemeral-storage and limits.ephemeral-storage with no additional changes.
Integration with Memory QoS
Memory QoS KEP-2570 configures cgroup v2 memory knobs at both container-level and pod-level cgroups to manage memory isolation and throttling as follows:
memory.min: Hard memory reclaim protection (configured for Guaranteed QoS pods), mapped from container or pod memory requests.memory.low: Soft memory reclaim protection (configured for Burstable QoS pods), mapped from container or pod memory requests.memory.high: Memory throttling threshold (configured for Burstable and BestEffort QoS pods at the container level). If a container’s memory usage crosses this threshold, the kernel reclaims memory aggressively and throttles all processes in that cgroup.memory.max: Hard memory limit (configured at both container and pod levels). If a cgroup’s memory usage reaches this limit and cannot be reduced, the kernel OOM killer is invoked. Memory QoS does not modify this knob; it remains mapped to standard container or pod memory limits.
Current Memory QOS settings
With KEP-2570, cgroup v2 knobs are calculated dynamically based on QoS classes and applied at both container-level and pod-level cgroups:
- Guaranteed QoS Pods:
- Container Level:
memory.min= container requestmemory.low,memory.high= disabledmemory.max= container limit
- Pod Level:
memory.min= sum of container requests (or pod-level request if specified)memory.low,memory.high= disabledmemory.max= sum of container limits (or pod-level limit if specified)
- Container Level:
- Burstable QoS Pods:
- Container Level:
memory.min= 0memory.low= container requestmemory.high=requests.memory + memory_throttling_factor * (limits.memory - requests.memory)limits.memorydefaults to node allocatable capacity if container limit is unset.
memory.max= container limit
- Pod Level:
memory.min= 0memory.low= sum of container requests (or pod-level request if specified)memory.high= disabledmemory.max= sum of container limits (or pod-level limit if specified)
- Container Level:
- BestEffort QoS Pods:
- Container Level:
memory.min,memory.low,memory.max= disabledmemory.high=memory_throttling_factor * node_allocatable_capacity
- Pod Level:
memory.min,memory.low,memory.high,memory.max= disabled
- Container Level:
Integration with DRA
Not including DRA allocations in memory cgroup settings triggers the following issues:
- If
memory.highis calculated based only on standard Spec limits, the container will suffer kernel reclaim at a threshold far below its actual allocated capacity. - If
memory.minormemory.lowis computed based strictly on standard Spec requests, the DRA memory allocation will be treated as unprotected, allowing the host kernel to reclaim it aggressively under system pressure.
Example Scenario (Without Integration)
Consider a Burstable container with a default memory throttling factor of 0.9:
- Container Spec:
requests.memory = 1GiB,limits.memory = 2GiB. - DRA allocation:
5GiBof direct memory. - Cgroup Configuration with Memory QoS:
memory.max=2GiB (Spec Limit) + 5GiB (DRA)=7GiB. (cgroup enforment section)memory.high=1GiB + 0.9 * (2GiB - 1GiB)=1.9GiB.- Outcome: Although the workload is allocated 7GiB of memory, its processes are actively throttled and compressed as soon as memory usage crosses 1.9GiB.
Memory QoS Settings with DRA
We maintain consistency with the CPU resource model. Kubelet applies a similar strategy for memory cgroups when a pod is allocated memory via a DRA ResourceClaim.
- Requests are inflated at the pod level and kept uninflated at the container level.
- Limits are inflated at both pod and container level.
- Container-Level Cgroups:
- Set container
memory.maxusing the inflated limit (limits.memory (Container) + DRA). - Set container
memory.min/memory.lowusing the uninflated Spec request. - Set container
memory.highusing the standard Memory QoS formula, but with the inflated limit (limits.memory (Container) + DRA) used formemory.maxcalculation.
- Set container
- Pod-Level Cgroups:
- Set pod
memory.maxusing sum of container limits + DRA, or pod-level limit if specified. - Set pod
memory.min/memory.lowusing sum of container requests + DRA, or pod-level request if specified. - Note:
memory.highis not set at the pod level with KEP-2570, so nothing changes here.
- Set pod
Example Scenario (With Integration)
Consider the same Burstable container under the integrated CPU-consistent configuration:
- Container Specification:
requests.memory = 1GiB,limits.memory = 2GiB. - DRA Memory claim allocation:
5GiBof direct memory. - Cgroup Configuration with Integration (CPU-Consistent):
memory.max=2GiB + 5GiB=7GiB.memory.high=1GiB + 0.9 * ((2GiB + 5GiB) - 1GiB)=6.4GiB.- Outcome: Throttling occurs correctly at 6.4GiB, allowing the container to utilize its full allocated 7GiB memory budget safely.
Pod Status Updates
Current Behavior:
- Allocated Resources (
pod.status.allocatedResourcesandpod.status.containerStatuses[*].allocatedResources):- Represents the desired intent or reservation. It publishes only requests.
- Kubelet sets this to match
pod.spec.containers[*].resources.requests(andpod.spec.resources.requestsat the pod level) upon successful pod admission or after successfully admitting a desired in-place resize.
- Resources (
pod.status.resourcesandpod.status.containerStatuses[*].resources):- Represents the actuated state or reality. It publishes both requests and limits.
pod.status.resources: Kubelet reads the actual requests and limits enforced on the pod-level cgroup directory directly from the host’s cgroup filesystem.pod.status.containerStatuses[*].resources: For running containers, Kubelet reads the cgroup state via CRI (e.g., CPU shares, quota, and memory limit).
Behavior with DRA: When DRA node allocatable resources are utilized, Kubelet enforces a split model to preserve intent tracking while accurately reporting actuated cgroup reality:
- Allocated Resources:
pod.status.allocatedResources: Set to pod-level resources if specified. If not, set to the sum of container-level standard requests and DRA requests.pod.status.containerStatuses[*].allocatedResources: No Change. It continues to be populated strictly based on standard requests in the PodSpec. For the Alpha scope, we do not plan to include container-level allocated resources to include DRA allocations as this field is not currently utilized for scheduler accounting. Shared claims across multiple containers make it difficult to attribute DRA resource allocation at the container status level. It continues to be populated strictly based on standard requests in the PodSpec.
- Resources (
pod.status.resourcesandpod.status.containerStatuses[*].resources):- Requests: Populated by reading the actual cgroup enforcement on the node. If the DRA driver/NRI plugin has adjusted these cgroup settings to actuate DRA resource allocations, the reported requests will reflect those changes. Since memory requests are currently not used to configure cgroup settings, we fallback to report what is requested in the spec and this would now include DRA requests.
- Limits: Populated by reading the actual cgroup enforcement on the node including DRA driver/NRI plugin modifications.
Kubelet Internal Resource States
In-place pod resizing and cgroup management introduce four distinct sets of resources that Kubelet tracks for each pod and container. The following defines these internal resource states and how they interact with DRA node allocatable resources:
- Desired Resources:
- What the user (or controller) asked for.
- Recorded in the API as the spec resources (
.spec.containers[i].resources). - Behavior with DRA: No change. Desired standard resources remain in
.spec.containers[i].resources, while DRA node allocatable resource claims are requested separately in.spec.resourceClaims.
- Allocated Resources:
- The resources that the Kubelet admitted, and intends to actuate.
- Persisted locally on the node in a checkpoint file.
- Used to update the pod status (
.status.allocatedResourcesand.status.containerStatuses[i].allocatedResources). - Behavior with DRA: No change. The node’s internal
allocatedcheckpoint remains strictly limited to standard Spec requests and limits at both the pod and container levels. DRA allocations are completely excluded.- Note: At the pod level, the API representation (
pod.status.allocatedResources) diverges from this internal state as the checkpoint does not include DRA requests. The pod status field accurately represents the total resource reservation including DRA, while the checkpoint remains spec-only to prevent Kubelet from triggering infinite resizing loops when comparing spec with the checkpointed state.
- Note: At the pod level, the API representation (
- Actuated Resources:
- The resource configuration that the Kubelet passed to the runtime to actuate.
- Not reported in the API.
- Persisted locally on the node in a checkpoint file.
- Behavior with DRA: No change. To ensure steady-state reconciliation loops (
computePodResizeAction) do not trigger unnecessary CRI updates or cgroup resets, Kubelet maintains the internalactuatedcheckpoint strictly limited to standard Spec requests and limits. DRA allocations are excluded from the checkpoint. - Divergence: This design introduces an intentional divergence where kubelet’s
actuatedcheckpoint excludes DRA allocations, diverging from the actual cgroup settings enforced on the node. This prevents the steady-state reconciliation loops from seeing a difference between.specand cgroups ensuring that we do not revert the DRA included cgroup settings.
- Actual Resources:
- The actual resource configuration the containers are running with, reported by the runtime, typically read directly from the cgroup configuration.
- Reported in the API via the
.status.containerStatuses[i].resourcesfield. - Behavior with DRA: During cgroup generation (
generateLinuxContainerResourcesandResourceConfigForPod), Kubelet dynamically inflates limits by summing standard Spec limits and DRA allocations read frompod.status.additionalNodeAllocatableResources. Therefore, the actual limits reported in.status.resources.limitsandcontainerStatuses[*].resources.limitsnatively reflect the combined standard and DRA resources based on the defined cgroup enforcement rules. Both initial pod creation and resize actuation share the exact same cgroup configuration code. Because these paths are identical, in-place vertical scaling preserves and applies the same DRA inflated cgroup values during actuation.
Integration with In-Place Pod Vertical Scaling
In the control plane, when the scheduler computes a resizing pod’s footprint, because PodRequests() aggregates
the DRA allocations from pod.status.additionalNodeAllocatableResources, the scheduler accurately tracks
total resource footprint during resize.
On the node, when Kubelet evaluates whether a resize fits on the node (canAdmitPod), the Allocation Manager
computes the resource footprint including DRA. When actuating the admitted resize at the container level we sum
the newly resized standard Spec limits with the constant DRA resources, passing the combined limits to CRI.
Kubelet Admission Control
The Kubelet has its own admission check
(AdmissionCheck)
to ensure a pod can run on the node, even after the scheduler has placed it. It utilizes the PodRequests() function from
the k8s.io/component-helpers/resource. This shared helper has been enhanced to support unified accounting. When
calculating a pod’s requirements, it aggregates the standard requests from pod Spec with the DRA allocations recorded in
pod.status.additionalNodeAllocatableResources. Because the scheduler populates this status field during the PreBind stage, the
Kubelet validates the pod’s comprehensive resource footprint.
This admission-time lookup reads directly from the pod.status.additionalNodeAllocatableResources API field. This allows Kubelet’s
Vertical Scaling Admission Controller (canAdmitPod inside AllocationManager) to accurately evaluate resource-fit during vertical resizing
without needing to persist DRA allocations in Kubelet’s local disk checkpoints (allocatedState or actuatedState). Because DRA allocations
are immutable after scheduling, Kubelet can bypass the local checkpoints for DRA evaluation, relying instead on this API status field as
the source of truth.
ResourceQuota Enforcement
ResourceQuota currently accounts for resources defined in the standard pod.spec requests/limits. A pod that receives CPU,
memory, or ephemeral storage through a DRA ResourceClaim consumes real node capacity, but no namespace quota tracks it.
This section describes how quota accounts for and enforces DRA node allocatable resources in Beta.
Two existing quota mechanisms remain unchanged:
- DRA device quota (
count/resourceclaims.resource.k8s.ioand<deviceclass>.deviceclass.resource.k8s.io/devices) charged from the claim spec at claim admission. - Standard spec quota (
requests.cpu/requests.memory) charged from the pod spec at pod creation.
To account for and enforce DRA node allocatable resources, the mechanism covers two operational cases:
Accounting for Running Pods with DRA Claims
When admitting a new pod that does not use DRA claims (or when the ResourceQuota controller calculates namespace
usage), quota must account for existing pods that consume DRA resources. Their footprints are persisted in pod.status.additionalNodeAllocatableResources.
As described in the scheduler and Kubelet sections, the shared component-helpers functions (PodRequests and
PodLimits) aggregate standard spec requests with DRA allocations from this status field. The core pod quota evaluator
(pkg/quota/v1/evaluator/core/pods.go) reuses these helpers to compute each pod’s total compute footprint. Both the
ResourceQuota admission plugin and the ResourceQuota controller share this evaluator, keeping ResourceQuota.Status.Used
in sync without requiring any separate accounting. The evaluator reads the status field without a feature gate check. The
field is only populated while the feature is enabled, and the unconditional read keeps kube-apiserver and
kube-controller-manager from disagreeing about usage when the gate is enabled on one but not the other.
Enforcement for Incoming Pods with DRA Claims
For standard spec resources, ResourceQuota admission runs synchronously at pod creation and rejects over-quota pods immediately.
When an incoming pod requests DRA node allocatable resources, create-time admission alone cannot prevent quota overcommitment because the claim’s
exact footprint depends on the selected node and its ResourceSlice definitions (such as deviceMultiplier or capacityMultiplier).
(Pods with explicit pod-level resources in pod.spec.resources are charged their full declared budget at creation, and the scheduler validates during
Filter that container requests plus the DRA footprint fit within that budget.)
For pods without pod-level resources covering the resource, quota is evaluated during PreBind when NodeResourcesFit.PreBind patches pod.status.additionalNodeAllocatableResources.
ResourceQuotaAdmission onpod/status: TheResourceQuotaadmission plugin inkube-apiservercurrently skipsstatussubresources. This would be updated to evaluatepod/statusupdates that modifyadditionalNodeAllocatableResourcesto a non-empty value.- On Quota Rejection: If the status patch exceeds namespace quota,
kube-apiserverrejects the patch.NodeResourcesFit.PreBindreturns an error andPreBindaborts beforeDynamicResources.PreBindruns (in the default plugin order), so no claim allocation is written to etcd. The scheduler callsUnreserve, clears the assumed pod from cache, and moves thePendingpod to the exponential backoff queue to retry until quota is freed or the pod is deleted. The non-spec additional node resources like DRA device capacity remain free for other pods to consume, while thepod.specresources already charged to quota at pod admission remain consumed, just like any other unscheduled pod in the scheduler queue. - On Quota Success:
ResourceQuotausage is updated in the API server, andDynamicResources.PreBindproceeds withbindClaimto allocate and reserve the claim before binding the pod. IfDynamicResources.PreBindorBindfails subsequently,NodeResourcesFit.Unreserveclearspod.status.additionalNodeAllocatableResources([]), releasing the quota.
Because the scheduler’s Filter stage does not consider quota, a pod in a namespace that is out of quota
still goes through node selection before evaluation in PreBind on each backoff retry (exponential backoff
bounds the retry churn). A better alternative is to evaluate quota at an earlier stage which is out of scope for
this KEP (see Scheduler-Side Quota Plugin).
HPA Integration
The HorizontalPodAutoscaler (HPA) controller compares a pod’s observed resource usage against its requested resources. Currently, HPA only reads requests from pod.spec and ignores pod.status.additionalNodeAllocatableResources, which has two implications for pods consuming node allocatable resources via DRA:
- Over-scaling: Because a pod’s actual CPU or memory usage includes resources consumed via its DRA allocation while HPA only considers the smaller
pod.specrequest, HPA computes an inflated utilization percentage and scales up excessively. - Unable to scale DRA-only pods: If a pod obtains all of its CPU or memory through a DRA
ResourceClaimwithout specifying a corresponding standard request inpod.spec, HPA treats the pod as having no request for that resource and fails to compute utilization at all.
To address this, the HPA controller will use the shared PodRequests helper when computing pod resource requests, incorporating both pod.spec requests and DRA allocations from pod.status.additionalNodeAllocatableResources for accurate scaling decisions.
Cluster Autoscaler Integration
Cluster Autoscaler decides scale-up and scale-down by simulating the scheduler over a snapshot of the cluster, and it measures node utilization by summing pod requests. Both paths read requests from the pod spec only, so DRA footprint is currently invisible to them.
Cluster Autoscaler already supports DRA. It snapshots ResourceClaims, ResourceSlices and DeviceClasses, and runs
the real scheduler plugins over that snapshot. When the DRANodeAllocatableResources feature gate is enabled, it will be updated to also account for DRA node allocatable resources:
- The autoscaler computes pod requests through the shared
component-helpersPodRequestsfunction so the additional node allocatable footprint is included. - The scheduler records the footprint in
pod.status.additionalNodeAllocatableResourcesduringPreBind, which the simulation does not run. When the feature gate is enabled and the scheduler plugin has computed the DRA footprint, the autoscaler fills in this status field the same way it fills in simulated claim reservations, and clears it when a pod is unscheduled.
Node Capacity Reporting
With DRA node allocatable resources, scheduling a pod with a node allocatable claim requires satisfying two constraints:
- Device capacity: Tracked by
ResourceSliceandResourceClaim. - Node allocatable capacity: Tracked by remaining
node.status.allocatable(node.status.allocatable- existing pod spec and DRA requests).
Because these two constraints are tracked in separate API objects, they are reported by different tools:
kubectl describe node: Will be updated to include DRA allocations in reported pod and node resource requests. However, this only reflects node resource headroom and does not show how many DRA devices or pool capacities remain unallocated inResourceSliceobjects.ResourcePoolStatusRequest(KEP-5677): Reports device-pool availability computed purely fromResourceSliceandResourceClaimstate. For pools whose devices declarenodeAllocatableResources, non-zero availability indicates free device entries in the pool, but the underlying node resources backing those devices may be consumed by regular pods without claims (since regular pod requests are node-scoped and do not specify which DRA pool or device supplies their capacity).
In summary, kubectl describe node provides overall node resource availability (accounting for both spec
and DRA requests), while ResourcePoolStatusRequest shows which specific DRA devices remain unallocated in
the pool.
Future Enhancements
Pass Allocation Details from Driver to Kubelet
Currently, Kubelet is blind to the type of resource allocation performed by the DRA driver. Passing this information from the DRA driver to Kubelet enables better
coordination and node-level cgroup enforcement. This can be solved by adding an AllocationType field inside NodeAllocatableResources and propagating
it all the way to the pod.status in the API.
The two types of allocation that can be configured are:
- Exclusive: Dedicates and physically isolates the resource capacity (e.g., pinning CPUs by setting
cpuset.cpusin the DRA driver). - Shared: Binds resources to a specific domain that is also shared with other containers not referencing the same claim (e.g., binding memory to a specific NUMA node, or binding CPUs to a socket that is also shared by other workloads).
API Changes
Device Spec (k8s.io/api/resource/v1/types.go):
// AllocationType specifies the isolation and scheduling strategy.
type AllocationType string
const (
// AllocationTypeShared indicates the resource is allocated from a shared general pool.
AllocationTypeShared AllocationType = "Shared"
// AllocationTypeExclusive indicates the resource represents dedicated, physically isolated capacity (e.g., dedicated cores).
AllocationTypeExclusive AllocationType = "Exclusive"
)
type Device struct {
// existing fields
// +optional
NodeAllocatableResources map[v1.ResourceName]NodeAllocatableResource
}
type NodeAllocatableResource struct {
Mapping *NodeAllocatableMapping
Overhead *NodeAllocatableOverhead
}
type NodeAllocatableMapping struct {
CapacityKey *QualifiedName
CapacityMultiplier *resource.Quantity
// AllocationType describes whether the resources represent exclusive or shared capacity.
// If omitted, it defaults to AllocationTypeShared.
// +optional
AllocationType *AllocationType `json:"allocationType,omitempty" protobuf:"bytes,3,opt,name=allocationType"`
}
Pod Status API (k8s.io/api/core/v1/types.go):
type PodStatus struct {
// ... existing fields ...
// +featureGate=DRANodeAllocatableResources
// +optional
AdditionalNodeAllocatableResources []AdditionalNodeAllocatableResource `json:"additionalNodeAllocatableResources,omitempty" protobuf:"bytes,25,rep,name=additionalNodeAllocatableResources"`
}
type AdditionalNodeAllocatableResource struct {
Source AdditionalNodeAllocatableReference
Containers []string
Mapping []NodeAllocatableMappedResources
Overhead []NodeAllocatableOverheadResources
}
type NodeAllocatableMappedResources struct {
Name ResourceName `json:"name" protobuf:"bytes,1,opt,name=name"`
Quantity resource.Quantity `json:"quantity" protobuf:"bytes,2,opt,name=quantity"`
// AllocationType is resolved from `device.nodeAllocatableResources.mapping.allocationType` and added here.
// +required
AllocationType AllocationType `json:"allocationType" protobuf:"bytes,3,opt,name=allocationType"`
}
Example:
ResourceSlice
apiVersion: resource.k8s.io/v1 kind: ResourceSlice metadata: name: native-resource-slice spec: driver: dra.cpu.com nodeName: my-node pool: { name: "node-pool", generation: 1, resourceSliceCount: 1 } devices: - name: socket0 attributes: "dra.example.com/type": "socket" allowMultipleAllocations: true capacity: "dra.example.com/cores": "64" nodeAllocatableResources: cpu: mapping: capacityKey: "dra.example.com/cores" capacityMultiplier: "2" allocationType: "Exclusive"Pod Status
"additionalNodeAllocatableResources": [ { "source": { "apiGroup": "resource.k8s.io", "kind": "ResourceClaim", "name": "cpu-claim" }, "containers": ["worker"], "mapping": [ { "name": "cpu", "quantity": "4", "allocationType": "Exclusive" } ] } ]
Node Cgroup Enforcement
Pod Level Cgroup
- CPU Limits: Set based on standard limits sum + unique direct mapped resources (refer to the Pod-Level Cgroup Settings calculation section above).
- CPU Requests:
- In the current alpha implementation, CPU shares are configured strictly based on the standard pod Spec requests sum.
- Under the proposed
AllocationType-aware future design:- Exclusive Mode (
AllocationType: Exclusive): Shares remain configured strictly based on the standard pod Spec requests sum. Since the DRA driver dedicates and physically isolates CPU capacity to the container (e.g., cpuset pinning), the workloads do not experience scheduling contention with other co-located pods on the node, making CFS shares inflation unnecessary. Setting shares based on exclusive resources reserved by the DRA driver also gives the container/pod an unfair advantage in the shared resource pool during resource contention. - Shared Mode (
AllocationType: Shared): Shares are inflated by adding the standard pod Spec requests sum and the resolved direct CPU quantity mapped by the claim (obtained frompod.status.additionalNodeAllocatableResources[].mapping[].quantity). Since the workload competes inside the node’s general shared resource pool, this inflation guarantees that the pod as a whole obtains its scheduler-reserved resources under contention.
- Exclusive Mode (
- Memory Limits: Set based on standard limits sum + unique direct mapped memory resources (refer to the Pod-Level Cgroup Settings calculation section above).
Container Level Cgroup
- CPU Limits: Set based on standard limits + direct resources + container overhead + pod overhead (refer to the Container-Level Cgroup Settings calculation section above).
- CPU Requests: Configured strictly based on the container’s standard Spec request (
pod.spec.containers[].resources.requests.cpu). - Why we do not set container-level shares based on DRA CPU:
- By setting the inflated CPU weight strictly at the pod-level parent cgroup, Kubelet guarantees correct resource priority relative to other pods in the cgroup hierarchy during resource contention. Inside the pod’s cgroup tree, sibling containers time-share the pod’s aggregate budget proportionally based on their relative standard Spec requests.
- When multiple containers in the same pod reference the same claim, dividing the claim’s CPU shares across container-level cgroups introduces complexity.
- Memory Limits: Set based on standard limits + direct resources + container overhead + pod overhead (refer to the Container-Level Cgroup Settings calculation section above).
- Memory Requests: Currently in kubelet, we do not set memory cgroups based on requests.
Enforcement Example:
# Pod Spec
spec:
containers:
- name: c1
resources:
requests: { cpu: "2", memory: "2Gi" }
limits: { cpu: "4", memory: "4Gi" }
claims: [{ name: "shared-claim" }]
- name: c2
resources:
requests: { cpu: "4", memory: "4Gi" }
limits: { cpu: "8", memory: "8Gi" }
claims: [{ name: "shared-claim" }]
resourceClaims:
- name: shared-claim
resourceClaimName: shared-cpu-claim
Depending on the allocation mapping type, cgroup parameters are actuated as follows:
1. Shared Mode (AllocationType = Shared)
In this case, the DRA driver allocates from a general node pool, so the status contains:
# Pod Status
status:
additionalNodeAllocatableResources:
- source:
apiGroup: resource.k8s.io
kind: ResourceClaim
name: shared-cpu-claim
containers: ["c1", "c2"]
mapping:
- name: cpu
quantity: "5"
allocationType: Shared
Cgroup bounds are set as:
- Pod Level Cgroup:
cpu.weight(CPU Shares): Inflated based on standard requests sum + DRA direct CPU quantity (2 + 4 + 5): 11 CPUs.cpu.max(CPU Quota): Set based on Pod-Level Cgroup Settings: 17 CPUs.memory.max(Memory Limit): Set based on Pod-Level Cgroup Settings: 12 GiB.
- Container Level C1 Cgroup:
cpu.weight(CPU Shares): Configured strictly based on standard container request: 2 CPUs.cpu.max(CPU Quota): Set based on Container-Level Cgroup Settings: 9 CPUs.memory.max(Memory Limit): Set based on Container-Level Cgroup Settings: 4 GiB.
- Container Level C2 Cgroup:
cpu.weight(CPU Shares): Configured strictly based on standard container request: 4 CPUs.cpu.max(CPU Quota): Set based on Container-Level Cgroup Settings: 13 CPUs.memory.max(Memory Limit): Set based on Container-Level Cgroup Settings: 8 GiB.
2. Exclusive Mode (AllocationType = Exclusive)
# Pod Status
status:
additionalNodeAllocatableResources:
- source:
apiGroup: resource.k8s.io
kind: ResourceClaim
name: shared-cpu-claim
containers: ["c1", "c2"]
mapping:
- name: cpu
quantity: "5"
allocationType: Exclusive
Cgroup bounds are set as:
- Pod Level Cgroup:
cpu.weight(CPU Shares): Kept uninflated, configured strictly based on standard requests sum (2 + 4): 6 CPUs.cpu.max(CPU Quota): Set based on Pod-Level Cgroup Settings: 17 CPUs.memory.max(Memory Limit): Set based on Pod-Level Cgroup Settings: 12 GiB.
- Container Level C1 Cgroup:
cpu.weight(CPU Shares): Configured strictly based on standard container request: 2 CPUs.cpu.max(CPU Quota): Set based on Container-Level Cgroup Settings: 9 CPUs.memory.max(Memory Limit): Set based on Container-Level Cgroup Settings: 4 GiB.
- Container Level C2 Cgroup:
cpu.weight(CPU Shares): Configured strictly based on standard container request: 4 CPUs.cpu.max(CPU Quota): Set based on Container-Level Cgroup Settings: 13 CPUs.memory.max(Memory Limit): Set based on Container-Level Cgroup Settings: 8 GiB.
Sharing Mapped Claims Across Pods
In the current design, ResourceClaims with direct node allocatable mapping entries can be shared across multiple containers within the same pod, but cannot be shared across different pods. Within a single pod, all containers share one parent pod cgroup that bounds their combined usage. Across multiple pods, however, each pod has its own separate pod cgroup, and pod.status currently has no way to indicate that a mapped resource is shared across pods—meaning PodRequests() would add the full mapped quantity to every pod, double-counting it against node capacity and inflating every pod’s cgroup request weights (cpu.weight).
However, this can be supported in a followup KEP by extending NodeAllocatableMapping in ResourceSlice and the Mapped API (DRANodeAllocatableMappedResources) in pod.status with AllowMultiplePods:
type NodeAllocatableMappedResources struct {
Name ResourceName
Quantity *resource.Quantity
// AllowMultiplePods indicates whether this mapped resource can be shared across multiple pods.
// +optional
AllowMultiplePods *bool
}
If allowMultiplePods is set to true:
- Scheduler and Admission: The scheduler, Kubelet admission, and Resource Quota account for the
mappedresource quantity only once. - Kubelet Cgroup Enforcement: Kubelet treats the
mappedquantity only as a limit. It skips adding themappedquantity to the pod’s cgroup requests (cpu.weight,memory.min/memory.low), and only adds it to the pod and container cgroup limits (cpu.max,memory.max,hugetlb.<pagesize>.max). This lets each pod sharing the claim burst up to the shared pool limit without over-reserving CPU shares or memory protection on the node, while the DRA driver enforces the shared pool boundary (for example, via a sharedcpuset.cpusmask or a size-bounded shared memory mount).
Scheduler-Side Quota Plugin
A dedicated scheduler side quota enforcement plugin could watch ResourceQuota objects and evaluate namespace quota during earlier scheduling phases (like Filter)
rather than relying on API server admission during PreBind. Rejecting pods over quota before PreBind allows the scheduler to return Unschedulable and hold the
pod in the unschedulable queue until a ResourceQuota update or pod deletion event occurs in the namespace, avoiding unnecessary node filtering and prebind.
Test Plan
[x] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.
Prerequisite testing updates
Unit tests
Unit tests will be added for all new and modified logic within the kube-scheduler and kubelet components.
- Ensuring the new fields in
DeviceandPodStatusare validated correctly, including the mapping combinations and the immutability ofadditionalNodeAllocatableResourcesonce the pod is bound. - Scheduler Plugin Logic (
NodeResourcesFit,DynamicResources):- Verify
NodeResourcesFitfilters on standard requests, includes the pod’s node-specific DRA allocations when they are already populated by DRA plugin. A node is feasible only if both plugin’s checks pass under either plugin order. - Verify the accurate calculation of a pod’s total node allocatable resource demand across both
Directmappings (device counts or capacity key drawdowns) andOverheadmappings (per-pod or per-reference auxiliary overheads). - Verify that inter-pod sharing of Mapped device claims is correctly blocked during the Filter stage, while inter-pod sharing of Overhead claims is permitted.
- Validating that
NodeResourcesFit.PreBindpatchespod.status.additionalNodeAllocatableResourcesbeforeDynamicResources.PreBind(bindClaim) runs, and thatNodeResourcesFit.Unreserveclears it if a subsequentPreBindorBindstep fails. - Verify the scheduler-assigned claim name for extended resources backed by DRA is recorded in the status patch and the claim is created with the same name only after the patch is accepted; a rejected patch creates nothing.
- Verify admin access results contribute no footprint and no status entry.
- Verify
- Scheduler Scoring (
NodeResourcesFit,NodeResourcesBalancedAllocation):- Verify scoring includes the DRA node allocatable footprint of the pod being scored and of existing pods on the node.
- Verify
NodeResourcesBalancedAllocationdoes not skip pods whose only CPU and memory come through claims.
- ResourceQuota (pod evaluator and quota admission):
- Verify namespace usage includes the DRA footprint (
mapping, per-pod and per-containeroverhead) and drops when the pod is deleted. - Verify a pod declaring pod-level resources produces a zero delta at the status write.
- Verify only writes changing
additionalNodeAllocatableResourcesare evaluated against quota; other status writes are unaffected.
- Verify namespace usage includes the DRA footprint (
- Scheduler Framework:
- Verify
NodeInfocache updates correctly in theAssumestage and reflects resources allocated to node allocatable resource claims. - Verify that when a pod using DRA node allocatable resources is deleted, the resources are correctly released and become available for other pods in the scheduler’s cache.
- Verify
- Component helper (
k8s.io/component-helpers/resource)- Testing the
PodRequestshelper function’s updated logic to include DRA node allocatable resources.- Ensure existing calculations for pods without DRA claims or PLR remain correct, properly aggregating init and regular container requests.
- Verify pod level resources when specified for a resource, continues to take precedence over per-container requests, include node allocatable claim requests.
- Verify that the node allocatable resources from
pod.status.additionalNodeAllocatableResourcesare correctly added to the pod’s effective standard resource requests. - Test that existing logic for different
PodResourcesOptions(e.g.,ExcludeOverhead,SkipPodLevelResources) continues to work as expected when DRA node allocatable resources are present, including correct handling ofpod.spec.overhead.
- Testing the
- Kubelet Admission Check
- Verifying that the admission check correctly uses the DRA node allocatable resource from the pod’s
status.additionalNodeAllocatableResourcesfield.
- Verifying that the admission check correctly uses the DRA node allocatable resource from the pod’s
- Kubelet Cgroup Enforcement (
pkg/kubelet/kuberuntime/kuberuntime_container_linux.go,pkg/kubelet/cm/helpers_linux.go):- Verify that container and pod-level CPU quota and memory limit correctly sum standard Spec limits and DRA allocations from
pod.status.additionalNodeAllocatableResources. - Verify that CPU shares remain purely based on standard Spec requests.
- Verify that container OOM score adjustments (
oom_score_adj) correctly incorporate DRA memory status allocations for Burstable pods using equal-splitting across referencing containers. - Verify cgroup generation across multiple test cases involving Pod Level Resources, including containers specifying their own limits and containers inheriting pod-level ceilings without limits.
- Verify that if container-level HugePages limits are omitted, Kubelet sets the limit to match DRA allocations.
- Verify that when
MemoryQoSis enabled, cgroup updates incorporate DRA memory allocations.
- Verify that container and pod-level CPU quota and memory limit correctly sum standard Spec limits and DRA allocations from
- Kubelet Allocation Manager (
pkg/kubelet/allocation/allocation_manager.go):- Verify that during steady-state reconciliation loops, Kubelet maintains the
allocatedcheckpoint strictly limited to standard Spec requests and limits, while correctly incorporating DRA allocations when evaluating node capacity during pod admission and resize checks.
- Verify that during steady-state reconciliation loops, Kubelet maintains the
- Kubelet Eviction and Preemption:
- Verify eviction ranking and preemption use requests that include the DRA footprint.
- Verify ephemeral storage: disk-pressure ranking, pod-level limit eviction, and container-level limit eviction account for DRA claim allocations; a pod without a spec limit stays unlimited; behavior is unchanged with the gate disabled.
- HPA request computation includes the DRA footprint when
pod.status.additionalNodeAllocatableResourcesis populated on the pod.
Coverage:
- pkg/scheduler/framework/plugins/dynamicresources: 20260916 - 87.0%
- pkg/scheduler/framework/plugins/noderesources: 20260916 - 91.3%
- pkg/scheduler: 20260916 - 83.2%
- pkg/scheduler/framework: 20260916 - 76.7%
- staging/src/k8s.io/component-helpers/resource: 20260916 - 99.2%
- pkg/kubelet/kuberuntime: 20260916 - 78.9%
- pkg/kubelet/cm: 20260916 - 28.6%
- pkg/kubelet/allocation: 20260916 - 85.7%
- pkg/kubelet/qos: 20260916 - 98.9%
- pkg/kubelet/status: 20260916 - 91.8%
- pkg/registry/core/pod: 20260916 - 82.8%
- pkg/apis/core/validation: 20260916 - 87.9%
- plugin/pkg/admission/nodedeclaredfeatures: 20260916 - 84.4%
- staging/src/k8s.io/component-helpers/nodedeclaredfeatures/features/dranodeallocatableresources: 20260916 - 100.0%
Integration tests
Integration tests are in test/integration/dra and cover the end-to-end scheduling flow.
The Kube-Scheduler filter/allocation and Kubelet items below were added as part of alpha. The ResourceQuota, Kube-Scheduler scoring, and HPA tests will be added with the beta implementation.
Kube-Scheduler:
- Tests to ensure correct interaction between
NodeResourcesFitandDynamicResourcesplugins. - Test that the scheduler’s internal cache (
NodeInfo.Requested) is accurately updated to reflect the resources consumed by pods with DRA node allocatable resource claims. - Ensure that resources are correctly released in the scheduler cache when a pod with DRA node allocatable resource claims is deleted.
- Validate that fungible claims resulting in different node allocatable resource footprints are accounted for correctly on a per-node basis.
- Verify that the scheduler correctly enforces inter-pod sharing restrictions, blocking pods that attempt to
share
Mappingdevices. - Tests to validate the
pod.status.additionalNodeAllocatableResourcesis populated correctly and the kubelet admission check correctly computes the effective pod resource request. - Test that
NodeResourcesFitscoring (LeastAllocated and MostAllocated) accounts for the DRA node allocatable footprint of both existing pods and the pod being scheduled across candidate nodes. - Test that
NodeResourcesBalancedAllocationscoring balances CPU and memory ratios incorporating DRA node allocatable requests.
Kubelet:
- Test that the Kubelet’s admission handler correctly factors in the node allocatable resources specified in
pod.status.additionalNodeAllocatableResourceswhen deciding whether to admit a pod. - Test that Kubelet correctly generates Linux cgroup configurations summing standard Spec limits and DRA allocations.
ResourceQuota:
- Test that when a new pod without DRA claims is created, namespace quota evaluation at admission includes the DRA footprints of existing scheduled pods.
- Test that a pod whose DRA node allocatable footprint fits the namespace budget binds successfully and that
ResourceQuota.Status.Usedreflects the standard requests plus the DRA footprint. - Test that a pod whose footprint exceeds the remaining namespace budget is rejected at
PreBind, remains unschedulable. - Test that deleting a bound pod releases the DRA portion of the usage.
- Test that a pod declaring pod-level resources covering its DRA footprint is rejected synchronously at creation when the namespace is out of quota.
HPA:
- Test that utilization calculation is DRA aware for pods with node allocatable claims.
Integration tests added for alpha:
e2e tests
E2E tests are added under test/e2e/dra and test/e2e_node.
Alpha:
- Verify these pods are scheduled onto nodes with sufficient capacity, considering both the pod’s standard requests and the DRA-added node allocatable resources.
These tests should cover various DRA modeling scenarios:
- Node allocatable resources as individual devices.
- Node allocatable resources as consumable capacity from a pool.
- Node allocatable resources from partitionable devices.
- Auxiliary node allocatable resources required by other devices (e.g., additional memory for an accelerator).
- Fungible claims involving node allocatable resources.
- Verify that Kubelet enforces correct cgroup settings and OOM score adjustments for CPU, memory, and HugePages.
Beta:
- Verify that a pod requesting DRA node allocatable resources exceeding namespace quota is rejected at
PreBind(NodeResourcesFitplugin) in a running cluster. - Verify Kubelet integration with Memory QoS and eviction under resource pressure.
- Verify HPA scaling based on pod resource requests that include DRA node allocatable resources.
e2e tests added for alpha: SIG Node, triage search
Graduation Criteria
Alpha
- Feature implemented behind the
DRANodeAllocatableResourcesfeature gate and disabled by default. - Core API changes for
DeviceandPodStatusintroduced. - Kube-Scheduler:
- The
DynamicResourcesplugin is updated to calculate and enforce node resource fit based on standard requests and node allocatable resource claims. - The scheduler’s internal cache update logic is enhanced to incorporate DRA node allocatable resource allocations.
- The
k8s.io/component-helpers/resourceshared library is enhanced to compute effective pod resource footprint.- The Kubelet’s admission handler is updated to consider node allocatable resource claims in
Pod.Status. - API validation restriction implemented in
pkg/apis/core/validation/validation.goblocking In-Place Pod Resizing for pods utilizing DRA node allocatable resources. - All unit and integration tests outlined in the Test Plan for the alpha scope are implemented and verified.
Alpha2
- Enhance Kubelet to utilize
pod.status.additionalNodeAllocatableResourcesfor cgroup management and OOM score adjustments. - Support use cases where DRA directly models node allocatable resources (such as exclusive CPU allocation or consumable capacity pools) as well as use cases where specialized devices declare auxiliary node allocatable resource dependencies (such as accelerator host memory overhead).
- Remove API validation restrictions in
pkg/apis/core/validation/validation.goto allow resizing standard Spec resources for pods utilizing DRA node allocatable resources. - Add E2E tests for kube-scheduler and Kubelet changes, including correct cgroup enforcement and OOM score adjustments across various device mapping models.
Beta
- One DRA driver (dra-driver-cpu, consumed by slurm-bridge) has integrated the API extensions and validated the node allocatable resource mapping in ResourceSlice.
- Kubelet eviction and preemption count DRA-allocated node allocatable resources, so pods are ranked by usage over their true request, not their spec request alone.
- Init container
perContaineroverhead counts only toward the pod’s peak resource calculation, the same way init container spec requests are counted. - Ephemeral storage - validation and the eviction manager integration.
- Unified
ResourceQuotaaccounting and enforcement for DRA-allocated node allocatable resources. - Compatibility with DRA-backed extended resources (scheduler-created claims) validated and tested.
- Per-node scoring in
NodeResourcesFitandNodeResourcesBalancedAllocation. - Requeue events for capacity: pods rejected on DRA-augmented capacity are requeued on pod deletion and pod scale-down, not only on claim and slice changes.
- HPA integration - pod utilization accounts for DRA node allocatable resources.
- Cluster Autoscaler integration - binpacking and node utilization account for DRA node allocatable resources, including for pods placed during simulation and pods duplicated into node templates.
kubectl describe nodeintegration - per-pod rows and the summary consider both spec and DRA values.- All unit, integration and e2e tests outlined in the Test Plan for the beta scope are implemented and verified.
- Feature gate
DRANodeAllocatableResourcesgraduates to Beta. Due to thev1.PodStatusfield rename in v1.38, it defaults tofalseunder the defaultMinCompatibilityVersion(N-1) andtruewhenMinCompatibilityVersionis1.38.
Beta 2
- Feature gate
DRANodeAllocatableResourcesis enabled by default (targeting v1.39) as the defaultMinCompatibilityVersionadvances to1.38. - Validation and feedback from DRA driver/s managing CPU and memory.
GA
- Validation and feedback from DRA driver/s managing CPU and memory.
- The feature has been enabled by default for at least two releases with no critical bug reports.
Upgrade / Downgrade Strategy
Upgrade: Enabling the feature gate on an existing cluster is safe. The new accounting logic will apply to any newly scheduled pods or pods that are re-scheduled. Existing pods with node allocatable resource claims would continue to run, but their claim request will not be reflected in the scheduler’s
NodeInfocache as these pods lackpod.status.additionalNodeAllocatableResourcesfield. On the node, Kubelet will continue to enforce cgroups based solely on standard Spec values for existing pods. To fully resynchronize control-plane accounting and node cgroup limit inflation, the pods with node allocatable resource claims must be restarted.pod.statusfield rename in Beta: Pods scheduled with the alpha feature gate enabled in v1.37 must be recreated after upgrading to v1.38:- In v1.38, the alpha
pod.status.nodeAllocatableResourceClaimStatusesfield is tombstoned and replaced bypod.status.additionalNodeAllocatableResources. - Running containers retain their existing DRA inflated cgroups during normal operation. However, because an upgraded v1.38 Kubelet ignores the tombstoned field, cgroups will revert to spec only values upon container restart, node reboot, or in-place resize.
- Recreating these pods allows the v1.38 scheduler to populate
pod.status.additionalNodeAllocatableResources, and restoring accounting and node cgroup enforcement.
- In v1.38, the alpha
Downgrade: Disabling the feature gate requires a kube-scheduler and kubelet restart. Upon startup, the scheduler rebuilds the NodeInfo cache without considering DRA node allocatable resources. The scheduler’s view of resource usage for existing pods will be incomplete (underestimated) as it does not consider claim-based requests, potentially leading to oversubscription of the node if new pods are scheduled. On the node, Kubelet will not dynamically trigger cgroup updates during regular sync loops. Running containers will continue to operate with their existing DRA-included cgroup limits. Kubelet will ignore
pod.status.additionalNodeAllocatableResourcesand revert cgroup limits to standard Spec only limits upon container restarts.
Version Skew Strategy
- API Skew: An older scheduler will not understand the new API fields. If
ResourceSliceorPodobjects contain the new fields, they will be ignored. - New Scheduler, Older Kubelet:
- To proactively prevent pods utilizing DRA node allocatable resources from landing on older Kubelets that do not enforce cgroup restriction based on DRA, the scheduler must use the Node Declared Features framework.
- A new declared feature
DRANodeAllocatableResourcesis registered innode.status.declaredFeatures. The API server admission controller uses this framework to validate and rejectResourceSliceobjects that containnodeAllocatableResourcesmappings if they target a node that lacks support for this feature.
pod.statusField Rename & Component Skew:pod.status.additionalNodeAllocatableResourcesin v1.38 replaces the alphanodeAllocatableResourceClaimStatusesfield. Because v1.37 components do not recognize the new field, theDRANodeAllocatableResourcesfeature gate usesMinCompatibilityVersionfor beta graduation:DRANodeAllocatableResources: { ... {Version: version.MustParse("1.38"), Default: false, PreRelease: featuregate.Beta}, {Version: version.MustParse("1.38"), Default: true, PreRelease: featuregate.Beta, MinCompatibilityVersion: version.MustParse("1.38")}, },- In v1.38, the feature gate remains
falseby default under the defaultMinCompatibilityVersion(1.37). - The feature can be enabled in v1.38 by setting
--feature-gates=DRANodeAllocatableResources=trueon all components, or by setting--min-compatibility-version=1.38. - In v1.39, it becomes enabled by default as the default
MinCompatibilityVersionadvances to1.38.
- In v1.38, the feature gate remains
- Pods scheduled under the v1.37 alpha gate must be recreated after upgrade (see Upgrade / Downgrade Strategy).
Production Readiness Review Questionnaire
Feature Enablement and Rollback
How can this feature be enabled / disabled in a live cluster?
- Feature gate (also fill in values in
kep.yaml)- Feature gate name:
DRANodeAllocatableResources - Components depending on the feature gate:
kube-scheduler,kubelet,kube-apiserver,kube-controller-manager.
- Feature gate name:
Does enabling the feature change any default behavior?
No. This feature only takes effect if users create Pods that request node allocatable resources via
pod.spec.resourceClaims and DRA drivers are installed and configured to expose node allocatable resources via
nodeAllocatableResources in ResourceSlice objects. Existing pods are unaffected.
Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?
Yes. Disabling the feature gate DRANodeAllocatableResources will prevent the scheduler from performing the unified accounting.
Pods already scheduled using DRA node allocatable resource accounting will continue to run. However, when new pods are scheduled
while the gate is disabled, any node allocatable resources specified in their DRA claims will not be considered by the scheduler.
This can lead to node oversubscription as the scheduler’s view of available resources on the node will be incomplete.
On nodes, running containers will continue to operate with their existing DRA included cgroup limits. Kubelet will only
ignore pod.status.additionalNodeAllocatableResources and revert cgroup limits back to standard Spec
limits upon subsequent container restarts or recreations.
What happens if we reenable the feature if it was previously rolled back?
The scheduler will resume its unified accounting logic for pods with DRA node allocatable resource claims. API
validation for the new fields will be re-enabled. The NodeInfo cache may be incorrect as it’s not
retroactively updated to consider node allocatable resource claims for previously scheduled pods. This inconsistent
state would persist until kube-scheduler restarts or all pods with node allocatable resource claims are restarted.
On nodes, running containers that were started while the gate was disabled will remain at standard Spec limits. To fully
resynchronize control-plane accounting and node cgroup limit inflation, pods utilizing DRA node allocatable claims must be restarted.
Are there any tests for feature enablement/disablement?
Feature enablement and disablement are tested through unit and integration tests.
Tests also verify that when the feature gate is disabled, existing pods with pod.status.additionalNodeAllocatableResources populated remain valid, while the field is dropped for new pods and ignored in resource calculations.
Rollout, Upgrade and Rollback Planning
How can a rollout or rollback fail? Can it impact already running workloads?
The scheduler prevents placing pods with DRA claims on older Kubelets that do not have the feature enabled using the Node Declared Features framework (DRANodeAllocatableResources).
Neither rollout nor rollback terminates, restarts, or evicts running pods. On rollback, running containers retain their DRA inflated cgroups. If a container restarts while the gate is disabled on Kubelet, its cgroups revert to standard Spec limits, which can cause CPU throttling or OOM kills if it relies on DRA capacity. Pods scheduled under the v1.37 alpha gate will have their status ignored by v1.38 Kubelets and must be recreated (see Upgrade / Downgrade Strategy).
What specific metrics should inform a rollback?
scheduler_unschedulable_pods(plugin="DynamicResources",plugin="NodeResourcesFit") orscheduler_pending_pods: A spike in unschedulable or pending pods indicating unexpected node capacity rejections orPreBindstatus patch failures.kubelet_admission_rejections_total: Pods rejected at Kubelet admission due to insufficient node allocatable capacity, indicating scheduler cache undercounting.container_cpu_cfs_throttled_seconds_totalor elevated container OOM kills: Workloads with DRA claims experiencing throttling or OOM kills due to cgroup limits not being inflated.ResourceQuota.Status.Used: Quota failing to decrease after pod deletion, or unexpected rapid depletion blocking standard workload admission.
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
Manual testing of the enable -> disable -> enable lifecycle was tested on Kind cluster with the feature gate enabled on 1.37:
- Upgrade (Enable): Enable feature gate on API server, scheduler, and Kubelet; deploy
dra-driver-cpuand a pod requesting DRA CPU. Verifypod.status.additionalNodeAllocatableResourcesis populated, scheduler cache reflects the combined footprint, and Kubelet inflates container cgroups. - Downgrade (Disable): Disable gate. Verify running pods continue uninterrupted and keep their recorded footprints in quota usage until deleted, new pods with DRA claims are scheduled without accounting and do not inflate cgroups, and restarting a container reverts its cgroup limits to standard Spec.
- Upgrade (Re-enable): Re-enable feature gate. Verify new pods regain unified accounting, pod status with
additionalNodeAllocatableResources, and cgroup inflation, and restarting earlier pods resynchronizes node cgroups.
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?
Yes. The alpha pod.status.nodeAllocatableResourceClaimStatuses field is replaced by pod.status.additionalNodeAllocatableResources in v1.38.
Pods using the feature with alpha feature gate enabled must be recreated after upgrade (see Upgrade / Downgrade Strategy).
Monitoring Requirements
How can an operator determine if the feature is in use by workloads?
ResourceSliceobjects containingDeviceentries withnodeAllocatableResources.- Pods with
status.additionalNodeAllocatableResourcespopulated.
How can someone using this feature know that it is working for their instance?
- API .status
- Other field:
pod.status.additionalNodeAllocatableResources - Details: Pods referencing node allocatable resource claims have this status field populated with their allocated footprint.
- Other field:
What are the reasonable SLOs (Service Level Objectives) for the enhancement?
Existing DRA and kube-scheduler SLOs continue to apply and must be maintained. Enabling this feature only affects pods requesting DRA Node Allocatable claims. Pods not using these claims should not experience any additional overhead. For pods utilizing DRA Node Allocatable claims, pod scheduling duration should be comparable to baseline DRA claim allocation and introduce no visible degradation compared to baseline scheduling performance.
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
- Metrics
- Metric names:
scheduler_plugin_execution_duration_seconds{plugin="DynamicResources", extension_point="Filter"}scheduler_plugin_execution_duration_seconds{plugin="NodeResourcesFit", extension_point="Filter"|"Score"|"PreBind"}scheduler_unschedulable_pods{plugin="DynamicResources"|"NodeResourcesFit"}scheduler_pending_podskubelet_admission_rejections_total
- Components exposing the metric:
kube-scheduler,kubelet
- Metric names:
Are there any missing metrics that would be useful to have to improve observability of this feature?
No.
Dependencies
Does this feature depend on any specific services running in the cluster?
No. The feature relies only on core Kubernetes components (kube-apiserver, kube-scheduler, kubelet) and compliant DRA drivers that publish ResourceSlice objects with node allocatable resource mappings.
Scalability
Will enabling / using this feature result in any new API calls?
Yes. The kube-scheduler issues a PATCH pods/status at PreBind for pods using node allocatable claims.
Will enabling / using this feature result in introducing new API types?
No. This KEP proposes extensions to an existing type, but not a new type itself.
Will enabling / using this feature result in any new calls to the cloud provider?
No.
Will enabling / using this feature result in increasing size or count of the existing API objects?
Yes. Individual ResourceSlice and Pod objects will have additional structured fields (nodeAllocatableResources
and additionalNodeAllocatableResources). However, because these fields are populated only for specialized workloads
utilizing DRA node allocatable claims, the overall cluster-wide memory and etcd storage footprint increase is minimal.
Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?
Yes. For pods utilizing DRA node allocatable claims, scheduling latency will slightly increase. The DynamicResources plugin
evaluates effective node capacity by summing standard Spec requests with DRA allocations. This increase is expected to be minimal.
Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?
No.
Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?
No
Troubleshooting
How does this feature react if the API server and/or etcd is unavailable?
Scheduling stops cluster-wide when the API server is unavailable, which is not specific to this feature. The PreBind status
patch fails, so the pod stays unbound and is retried.
Already-bound pods are unaffected. Kubelet keeps enforcing cgroups from the last observed pod status, and running containers are not affected. Quota enforcement runs inside the API server, so it is also unavailable when the API server is down.
What are other known failure modes?
[Pod unschedulable due to unresolvable device mapping during driver slice churn]
- Detection:
scheduler_unschedulable_pods{plugin="DynamicResources"}rises; podFailedSchedulingevents report that device mappings could not be resolved. - Mitigations: Recovers automatically when the DRA driver republishes its
ResourceSlices. Restart the driver daemonset if slices remain absent. - Diagnostics: Scheduler logs at
-v=5show device mapping lookup failure inDynamicResourcesplugin during fail-closed sharing validation. - Testing: Scheduler unit tests for fail-closed sharing validation.
- Detection:
[Pod pending at PreBind due to namespace ResourceQuota exhaustion]
- Detection: Pod status remains
Pending;scheduler_pending_pods{queue="backoff"}andscheduler_plugin_execution_duration_seconds{plugin="NodeResourcesFit", extension_point="PreBind", status="Error"}rise; pod events reportFailedSchedulingwithexceeded quota. - Mitigations: Increase the namespace
ResourceQuotaor reduce claim requested quantities. - Diagnostics: Scheduler logs at
-v=5show the quota rejection error on the status patch inNodeResourcesFit. - Testing: Integration tests covering
PreBindquota rejection.
- Detection: Pod status remains
What steps should be taken if SLOs are not being met to determine the problem?
N/A. This feature does not define any separate SLO; general kube-scheduler and DRA SLOs apply.
Implementation History
- 2025-12-22: KEP created.
- v1.36: Initial alpha. Device API extensions, scheduler fit checks, and pod status recording.
- v1.37: Alpha2. Kubelet cgroup and OOM enforcement.
Drawbacks
Alternatives
DeviceClass API Extension for NodeAllocatableResourceMappings
In this option, the primary information about how a DeviceClass relates to node allocatable resources is contained within the DeviceClassSpec.
// In k8s.io/api/resource/v1/types.go
type DeviceClassSpec struct {
// ... existing fields
// NodeAllocatableResourceMappings lists the node allocatable resources that this DeviceClass can provide or depend on.
// +optional
// +featureGate=DRANodeAllocatableResources
NodeAllocatableResourceMappings []NodeAllocatableResourceMapping `json:"nodeAllocatableResourceMappings,omitempty"`
}
// NodeAllocatableResourceAccountingPolicy, NodeAllocatableResourceQuantity
// are defined the same as in the main proposal.
Reason for Not Choosing:
While defining NodeAllocatableResourceMappings in the DeviceClass is simpler, it lacks the granularity needed for many real-world scenarios. The Device API Extension approach allows these mappings to be specified per-Device instance within the ResourceSlice. This is advantageous because:
- Heterogeneous Devices: Even within the same
DeviceClass, individual device instances can have different node allocatable resource implications. For example, different GPU models or even the same model on different parts of the system topology might have varying CPU/memory overheads. Option 1 cannot express this. - Complex Resources: Resources where we use Partitionable Devices to model hierarchies (e.g., sockets, NUMA nodes, caches, cores). The node allocatable resource capacity (e.g., number of CPUs) is associated with specific instances in the hierarchy changes and this is best represented in individual
Deviceentries.
Explicit AccountingPolicy in DeviceClass and PodStatus
In the initial Alpha 1 proposal (KEP_orig.md), future enhancements for accounting policies explored defining an explicit string enum NodeAllocatableResourceAccountingPolicy configured inside DeviceClass and tracked in PodStatus.
// NodeAllocatableResourceAccountingPolicy defines how node allocatable resource quantities like CPU, Memory
// allocated via DRA are aggregated with standard resource requests in the PodSpec.
type NodeAllocatableResourceAccountingPolicy string
const (
// PolicyAddPerClaim indicates that the node allocatable resource quantity in the DRA claim
// is treated as additional to the pod spec requests. This quantity is accounted
// for exactly once per claim instance, regardless of the number of containers referencing it.
PolicyAddPerClaim NodeAllocatableResourceAccountingPolicy = "AddPerClaim"
// PolicyAddPerReference indicates that the node allocatable resource quantity in the DRA
// claim is treated as additional to the pod spec requests. This quantity is
// accounted for cumulatively for every reference to the claim.
PolicyAddPerReference NodeAllocatableResourceAccountingPolicy = "AddPerReference"
// PolicyMax indicates that effective request is the greater value between the standard container
// request and the DRA claim for the same resource.
PolicyMax NodeAllocatableResourceAccountingPolicy = "Max"
// PolicyConsumeFrom indicates that a DRA claim is defined to represent the node
// resource pool capacity. All containers or pods referencing the claim are satisfied from the capacity pool defined by the DRA claim.
PolicyConsumeFrom NodeAllocatableResourceAccountingPolicy = "ConsumeFrom"
)
// In k8s.io/api/resource/v1/types.go
type DeviceClassSpec struct {
// ... existing fields ...
// NodeAllocatableResourceAccountingPolicies defines how the node allocatable resource represented by the devices
// in this class should be accounted for and aggregated with any standard request for the same resource.
// +optional
// +featureGate=DRANodeAllocatableResources
NodeAllocatableResourceAccountingPolicies map[ResourceName]NodeAllocatableResourceAccountingPolicy
}
// In k8s.io/api/core/v1/types.go
type AdditionalNodeAllocatableResource struct {
// ... existing fields ...
// AccountingPolicy tells Kubelet which policy was used by the scheduler.
AccountingPolicy map[ResourceName]NodeAllocatableResourceAccountingPolicy
}
Reason for Not Choosing:
- Granularity for Additive Policies: While configuring an explicit
AccountingPolicyenum inDeviceClassis simpler, it lacks the granularity needed for complex device configurations. Instead, Alpha 2 transitions to a structured union directly insideDevice.NodeAllocatableResourceMappings(DirectvsOverhead). This allows drivers to declare exact capacity consumption or auxiliary per-pod/per-reference overheads natively per device instance rather than forcing a single flat policy across an entire device class. - Generic Reservation Solution for ConsumeFrom: The
ConsumeFrompolicy (reserving a pool of resources and drawing down from it) represents a much broader concept that applies to all cluster resources, not just node allocatable resources. Attempting to solveConsumeFromexclusively within this KEP would create duplicate, domain-specific reservation mechanisms. Therefore,ConsumeFromis excluded from this KEP in favor of a unified, generic reservation solution being explored in Kubernetes Enhancement Issue #6048.
Alternative Model for pod level resources + DRA
The current model treats pod-level resources as the upper bound: the declared pod-level values are the pod’s effective footprint, and DRA-delivered resources must fit within them. The alternative treats them as independent: the pod-level values cover the containers, and the DRA footprint is added on top. When pod-level resources are not specified, the two models are identical: the footprint is the sum of container requests plus DRA allocations.
Ceiling (current): pod-level resources include the DRA footprint.
- Pros:
- Pod-level resources keep representing the whole pod footprint.
- Claims are requested at the container level, so adding the claim’s resources to the container’s spec requests matches the spec shape, and the pod-level value stays the bound over both.
- Existing DRA drivers (like dra-driver-cpu) and their users can keep accounting correct while the feature rolls out by declaring pod-level bounds that include the claim.
- Quota and autoscaler integrations keep working: the pod-level value is the effective footprint they already read.
- Cons:
- Does not work well with pod-level resources plus prioritized lists: the footprint depends on the selected alternative, so users must over-declare or omit pod-level resources.
- A driver upgrade that raises overhead can make existing pods unschedulable when the footprint no longer fits the pod-level budget.
Additive: the DRA footprint is added on top of pod-level resources.
- Pros:
- Node-variable overhead and prioritized lists work without over-declaring.
- A driver upgrade does not invalidate existing pod specs.
- Cons:
- There is no way to specify per-pod total cap.
Future option: both models are implementable at the scheduler and for node enforcement, so a configuration option to pick between the two can be added if a use case emerges.