KEP-5690: DRA Preemption
KEP-5690: DRA Preemption
- Release Signoff Checklist
- Summary
- Motivation
- Proposal
- Design Details
- Production Readiness Review Questionnaire
- Implementation History
- Drawbacks
- Alternatives
- Infrastructure Needed (Optional)
Release Signoff Checklist
Items marked with (R) are required prior to targeting to a milestone / release.
- (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
- (R) KEP approvers have approved the KEP status as
implementable - (R) Design details are appropriately documented
- (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
- e2e Tests for all Beta API Operations (endpoints)
- (R) Ensure GA e2e tests meet requirements for Conformance Tests
- (R) Minimum Two Week Window for GA e2e tests to prove flake free
- (R) Graduation criteria is in place
- (R) all GA Endpoints must be hit by Conformance Tests within one minor version of promotion to GA
- (R) Production readiness review completed
- (R) Production readiness review approved
- “Implementation History” section is up-to-date for milestone
- User-facing documentation has been created in kubernetes/website , for publication to kubernetes.io
- Supporting documentation—e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes
Summary
Dynamic Resource Allocation (DRA) was promoted to stable in v1.35, but one major capability—preemption of pods utilizing DRA resources—is currently not supported. Under the current state, lower-priority workloads that happen to be scheduled early can occupy scarce devices indefinitely, leaving higher-priority workloads pending. By introducing native preemption support in DRA, we equip the scheduler with the necessary capabilities to ensure that the most critical workloads gain access to scarce hardware resources.
Simulating the removal of victims is not sufficient on its own. The devices held by preempted pods are reclaimed asynchronously by the resourceclaim controller, so there is a window during which a preemptor has been promised capacity that it does not yet hold. This KEP therefore also specifies how the scheduler records and holds that capacity, so that a preemptor is not made to preempt repeatedly and its promised resources are not taken by other pods while being reclaimed.
Most of this is contained in the dynamicresources plugin. The one exception is a new optional
scheduler framework extension point, NominationExtensions, which notifies plugins when a
preemption candidate is selected or cleared and lets a plugin report whether resources freed by an
earlier preemption are still being reclaimed.
Motivation
Acquisition and allocation of specialized hardware (e.g., GPUs and TPUs) is a primary operational concern for AI/ML users. Because accelerator resources are frequently constrained, it is essential that Kubernetes provides robust tools to run high-priority workloads on the available hardware. Introducing preemption support for DRA ensures that the scheduler can reclaim resources from lower-priority workloads to satisfy scheduling requirements of higher-priority tasks.
Goals
- Enable the scheduler to preempt pods consuming node-local DRA devices, in the pod-by-pod
preemption path implemented by
DefaultPreemption. - Ensure that a preemptor is not made to preempt repeatedly because the devices it freed are reclaimed asynchronously, or are taken by another pod before it can be scheduled.
Non-Goals
- Support preemption in the workload-aware preemption path. Pods belonging to a PodGroup are out of scope, whether their ResourceClaims are reserved for the PodGroup or for individual pods. What this would require is outlined in Deferred: workload-aware preemption .
- Support preemption of pods using multi-node or network-attached devices.
- Persist the scheduler’s record of in-flight preemptions across a scheduler restart or expose it in the API for external components such as Cluster Autoscaler.
- Coordinate held capacity between multiple schedulers.
Proposal
We will implement the fwk.PreFilterExtensions interface in the dynamicresources scheduler plugin, which
requires implementing the AddPod and RemovePod methods. These functions are invoked by the core
DefaultPreemption plugin to incrementally simulate the removal or recovery (reprieval) of candidate victim
pods during the preemption planning loop, as well as by the scheduling framework when accounting for
nominated pods on a node (RunFilterPluginsWithNominatedPods).
Implementing these functions requires updating the internal stateData structure in the dynamicresources
plugin to track state transitions and maintain local capacity maps transactionally throughout the preemption
simulation.
The scheduler must correctly handle several advanced DRA features during this simulation, or explicitly
bypass preemption if those features are in use.
Features that we will support but require careful implementation:
- Partitionable Devices: Freeing up the exact device partition requested by a higher-priority workload may require preempting multiple lower-priority pods that collectively occupy fractional partitions (represented as counters) on the same physical device, so that the raw capacity can be consolidated and re-partitioned.
- Consumable Capacity: We might need to free up just a subset of the capacity on a device to satisfy the request of a higher-priority pod. We need to make sure the remaining capacity on a device is correctly tracked during preemption simulations, and that we don’t overcommit.
- Device Binding Conditions: A lower-priority pod may be blocked in
PreBindwaiting for device binding conditions to be satisfied while already holding an in-flight or allocated ResourceClaim. Because the pod has already been assumed inReserve, its allocation is tracked indraManager(and must be released byRemovePodduring preemption simulation), and the scheduler’s preemption executor already cancels thePreBindcontext when the pod is preempted so the claim can be deallocated.
Features/scenarios that we will not support:
- ResourceClaims that span multiple nodes, and network-attached devices: Every DRA device has an
associated node selector (
nodeName,nodeSelector, orallNodes), either set directly on the device or inherited from theResourceSlicein which the device is defined. Thedynamicresourcesplugin will only simulate the release of claims whose allocated devices all use the explicitnodeNameselector (Spec.NodeName != nil). Claims usingnodeSelectororallNodesare ignored during preemption simulation. Separately, preemption enumerates candidate victims from the pods on the node being evaluated, so a device held by a pod on a different node is never offered as something that could be freed. - Pods belonging to a PodGroup: Preemption for these goes through the workload-aware preemption path, which removes victims from the simulation differently. The plugin will not free devices for them, whether their ResourceClaims are reserved for the PodGroup or for individual pods. See Deferred: workload-aware preemption .
- Extended resources backed by DRA (
DRAExtendedResource): For pods requesting extended resources backed by DRA, the backingResourceClaimis synthesized in memory duringFilterand is only created in the API server duringPreBind(and recorded inpod.Status.ExtendedResourceClaimStatusrather thanpod.Spec.ResourceClaims). Preemption of or by pods using DRA-backed extended resources is not supported in Alpha.
Simulating the removal of victims is not the whole problem. The devices freed by a preemption do not
become allocatable when the victim pods are deleted, but later, when the resourceclaim controller
deallocates their ResourceClaims. During that interval the scheduler can preempt further pods
unnecessarily, and the freed devices can be taken by an unrelated pod. When a preemption candidate is
selected, dynamicresources records in memory the simulated AllocationResults computed for the preemptor
together with the victim ResourceClaims being released. It uses the simulated AllocationResults
in AddPod to hold the exact capacity the preemptor needs on the nominated node, and uses the victim
ResourceClaims to defer any further preemption by the same pod until those claims have been
deallocated. This is specified in Design Details
.
Supporting this requires a small, generic addition to the scheduling framework: an optional
NominationExtensions interface. Because DefaultPreemption evaluates candidate nodes concurrently
using cloned CycleStates that are discarded when DryRunPreemption ends, a plugin cannot tell
from RemovePod and Filter alone which candidate won. NominationExtensions notifies the plugin
when a winning preemption candidate is selected (passing the scheduling cycle’s CycleState, the
winning node, and the victim pods) or when a nomination is cleared, and lets a plugin
report in PodEligibleToPreemptOthers whether resources freed by an earlier preemption by that pod
are still being reclaimed. Today DefaultPreemption answers that eligibility question with a
hard-coded heuristic: if a pod already has nominatedNodeName set from an earlier preemption, it
refuses to let that same pod preempt again while any pod on its nominated node is still terminating
(other incoming pods without nominatedNodeName set are still free to preempt on that node). That
heuristic works for resources released with the pod, but not for resources that a controller
reclaims afterwards. NominationExtensions makes both the nomination lifecycle and the settling
check pluggable. For Alpha, NominationExtensions is provisional (and may be kept internal to the
framework rather than exposed to external plugins) while SIG Scheduling evaluates whether DRA state
and nominations should move into the core scheduler framework for Beta.
Risks and Mitigations
The risks below concern the claim nomination mechanism specified in Claim nomination , which holds the capacity promised to a preemptor until the preemptor has been scheduled. The timeline referred to as t0 to t4 is defined in Preemption settling for DRA .
- Nominations are in-memory and not visible to external components or across scheduler restarts.
Nominations are held in memory in the
dynamicresourcesplugin. A scheduler restart during the settling window returns the affected pods to the behavior they would have without this feature: the preemptor may have its freed devices taken by another pod, and may cascade once. Similarly, external components that run scheduling simulations—most notably Cluster Autoscaler and external queue controllers such as Kueue—can seenominatedNodeNameon the preemptor pod in the API, but cannot see which specific DRA devices or capacities on that node have been nominated for it, and may therefore make conflicting scale-down or placement decisions during the settling window. Persisting nominations in the API server (for example inResourceClaim.Status, or on the Pod for cases such asDRAExtendedResourcewhere noResourceClaimobject exists beforePreBind) requires an API change; we defer that design to Beta. - Claims that are never deallocated. If the resourceclaim controller is unhealthy and fails to
deallocate a victim claim, its capacity remains occupied in the API server (which is standard DRA
behavior whenever a pod is deleted while the controller is down). The preemptor remains waiting
for its nominated node—matching how
DefaultPreemptionbehaves when a victim pod is stuck terminating—rather than timing out and evicting additional victims on other nodes while the controller is unhealthy. Once the controller recovers and deallocates the claim (or if the preemptor is deleted), the nomination resolves. - Multiple schedulers. A nomination is only known to the scheduler that created it. Another scheduler may allocate the held capacity. This matches the existing limitation of nominated nodes.
Design Details
Preemption simulation
During PreFilter, dynamicresources snapshots AllocatedState from draManager into stateData in
CycleState. During DryRunPreemption, DefaultPreemption clones CycleState (stateData.Clone()) for
each candidate node and drops it when the simulation ends. RemovePod and AddPod only read from
draManager and mutate the cloned AllocatedState in CycleState (which Filter uses to construct a
local allocator), so draManager itself never needs to be mutated or cloned.
Preemption settling for DRA
When the scheduler decides to preempt, it deletes the victim pods, sets nominatedNodeName on
the preemptor and returns the preemptor to the scheduling queue. There is then a window during
which the preemptor has been promised resources that it does not yet hold:
| Time | Event | Actor |
|---|---|---|
| t0 | Victim pods are deleted and graceful termination begins | kube-scheduler |
| t1 | The victim pod objects are gone | kubelet / API server |
| t2 | ReservedFor is cleared on the victims’ ResourceClaims | resourceclaim controller |
| t3 | Status.Allocation is cleared; the devices become allocatable again | resourceclaim controller |
| t4 | The preemptor is retried and scheduled | kube-scheduler |
The scheduler already handles part of this window. When a pod that already has nominatedNodeName
set is retried, PodEligibleToPreemptOthers refuses to let that pod start a second preemption while
its nominated node still holds pods that are terminating because of preemption (while still allowing
other un-nominated pods to preempt on that node). That check assumes the settling window ends when
the victim pods are gone, that is at t1.
For DRA the assumption does not hold, because devices are reclaimed asynchronously by the resourceclaim controller and only become allocatable at t3. Two problems follow:
- Cascading preemption. Between t1 and t3 the terminating-pod check no longer applies, but the devices are still allocated. If the preemptor is retried during this interval it fails again and preempts a second set of victims that were never needed. In the worst case, a slow or stuck resourceclaim controller turns a single preemption into a series of them.
- Device stealing. Between t3 and t4 the devices are free, and nothing in the API records that they were freed on behalf of the preemptor. An unrelated pod may be allocated them first, after which the preemptor has to preempt all over again. Under contention for scarce accelerators this can repeat indefinitely, which would defeat the purpose of this KEP.
These are two views of the same gap: the scheduler has no representation of capacity that a preemptor has earned by preempting but has not yet received. A single mechanism, described below, addresses both.
DRA is not fundamentally special here. The same gap exists for any resource that is reclaimed by a controller rather than by the removal of the pod itself; DRA is simply the first case where the interval is long enough to matter in practice.
Claim nomination
When DefaultPreemption selects a winning candidate for a preemptor pod, it invokes
NominationExtensions.AddNominatedPod on registered plugins. Because CycleState only exists for a
single scheduling cycle, dynamicresources records two pieces of cross-cycle state in draManager
(analogous to PodNominator in the scheduler cache):
- Per preemptor pod (claim nomination): the nominated node and the simulated
AllocationResultfor each of the preemptor’sResourceClaims on that node; - Per node (
nodeName): the UIDs of the victimResourceClaims on that node that the simulation released and that are waiting for deallocation (Status.Allocation != nil). Conceptually, deallocating claims are tracked at the scope of each claim’sStatus.Allocation.NodeSelector; in Alpha, where only node-local devices are supported, this is always a singlenodeName.
When DefaultPreemption actuates the winning preemption candidate in PostFilter, it invokes
NominationExtensions.AddNominatedPod with the scheduling cycle’s CycleState, the winning
nodeName, and victims. In AddNominatedPod, dynamicresources derives the released victim
ResourceClaim UIDs from victims (including only claims whose every reserving pod is in victims
or already terminating, not shared claims still held by a non-preempted pod) and records them for
nodeName in draManager. It then constructs the allocated device state with victims removed
(and any higher/equal-priority nominated pods on nodeName applied) and runs the deterministic DRA
allocator for nodeName to compute and store the preemptor’s simulated AllocationResults in
draManager. In subsequent scheduling cycles, PreFilterExtensions.AddPod looks up the nominated
pod’s AllocationResults from draManager and simulates allocating them in the current cycle’s
CycleState. Likewise, when the preemptor pod itself is evaluated on its nominatedNodeName during
regular scheduling (evaluateNominatedNode), dynamicresources.Filter first checks whether its
nominated AllocationResults in draManager are feasible against the current allocated device state
and uses them directly, falling back to a full allocator search on the node only if those nominated
AllocationResults are no longer feasible.
A nomination’s lifetime is bound 1-to-1 to the pod’s nominatedNodeName in the scheduler’s
PodNominator. It is created by NominationExtensions.AddNominatedPod and discarded by
NominationExtensions.RemoveNominatedPod when:
- the preemptor is scheduled and bound;
- the preemptor is deleted;
- the preemptor’s
nominatedNodeNameis cleared or replaced by a subsequent preemption.
Once victim pod deletion has been actuated, the victim ResourceClaim UIDs recorded for nodeName
remain tracked on nodeName in draManager until their Status.Allocation is cleared at t3, even
if the original preemptor’s nomination is later removed or replaced. No separate time-based expiry is
needed. Once all victim claims on nodeName have been deallocated at t3, if the
preemptor is retried and still cannot be scheduled on the nominated node (for example because a
higher-priority pod took the freed device, or the node became unschedulable), PodEligibleToPreempt
returns true and DefaultPreemption immediately re-evaluates the pod—either replacing the
nomination with a new candidate node or clearing nominatedNodeName, which removes the nomination.
Likewise, if the API call to delete a victim pod fails during preemption execution, the scheduler’s
preemption executor immediately clears the nomination (RemoveNominatedPod), discarding the recorded
victim claims so preemption can be retried without waiting for a timeout.
Before t3, while the victim claims are still allocated, waiting without a timeout matches how
DefaultPreemption behaves when a victim pod is stuck terminating: if the resourceclaim controller
is unhealthy, expiring the nomination would only cause the preemptor to evict additional victims on
other nodes while the controller is down.
Because a nomination is held only in memory, removing it when a preemptor is deleted or its
nominatedNodeName is cleared or changed does not produce a ResourceClaim event in the API server.
Just as the scheduler emits a synthetic EventAssignedPodDelete when a pod nomination is removed to
mimic an assigned pod deletion, RemoveNominatedPod emits a synthetic ResourceClaim update event
(mimicking deallocation of the nominated AllocationResults) to the scheduling queue when the node’s
tracked victim claims already have Status.Allocation == nil (after t3), allowing dynamicresources’s
existing ResourceClaim QueueingHint (isSchedulableAfterClaimChange) to wake unschedulable pods
that were blocked by the held capacity. Before t3, while the victim claims still have
Status.Allocation != nil, no synthetic event is needed because the real ResourceClaim
deallocation event at t3 will wake those pods.
Nominations are held in memory in the dynamicresources plugin and are keyed by the preemptor’s pod
UID. They are not persisted; see Risks and Mitigations
.
A claim nomination is distinct from, but paired with, the pod’s nominatedNodeName. The latter is an
API field recording which node the scheduler intends to place the pod on; the former is scheduler-local
state recording which DRA capacity on that node is being held for the preemptor (alongside the
per-node record of victim claims being deallocated on that node). A claim nomination never exists
without a corresponding nominatedNodeName.
Together, the preemptor’s claim nomination and the node’s tracked victim claims serve two purposes: they hold the promised capacity for the preemptor, and they defer any further preemption by a pod nominated on that node while its victim claims are still being reclaimed. The subsections below cover the first, then how it composes with the preemption simulation, then the second.
Holding capacity for the preemptor
When the scheduler evaluates a node for another pod Q—either during normal Filter or during a
preemption simulation—the framework’s RunFilterPluginsWithNominatedPods evaluates Filter in up
to two passes whenever nominated pods P with priority >= Q exist on that node:
- Pass 1 (with nominated pods):
RunFilterPluginsWithNominatedPodsappliesPreFilterExtensions.AddPodfor each such nominated podPand runsFilter. - Pass 2 (without nominated pods): If Pass 1 succeeds,
RunFilterPluginsWithNominatedPodsrunsFiltera second time withoutAddPod(P)applied.
When AddPod is called for a nominated preemptor P in Pass 1, dynamicresources consults P’s
live nomination and performs two updates to the allocated device state in CycleState:
- Subtract still-allocated victim claims (t0 to t3): For each victim
ResourceClaimUID on that node that still hasStatus.Allocation != nilin the informer cache (and has not already been removed inCycleState),dynamicresourcessubtracts itsStatus.Allocationfrom the allocated device state, using the same accounting helper asRemovePod. This prevents double-counting the capacity thatPitself is taking from its victim claims between t0 and t3: in particular, between t1 and t3, the victim Pod object has already been deleted from the API server and is no longer inNodeInfo.Pods(soRemovePodis not called for it), yet itsResourceClaim.Status.Allocationremains present in the state built byPreFilteruntil t3. - Add the preemptor’s simulated
AllocationResults (t0 to t4):dynamicresourcesaddsP’s recordedAllocationResults into the allocated device state, marking the exact devices, consumable capacity shares, or partition counters selected forPas in use.
When both Pass 1 and Pass 2 run, Pass 2 of dynamicresources.Filter verifies that the exact
AllocationResult computed for Q in Pass 1 (which respected AddPod(P) and avoided P’s
nominated devices) is also valid against the un-nominated state in Pass 2, rather than computing a
different AllocationResult in Pass 2. This ensures a single AllocationResult is valid in both
passes before saving it to nodeAllocations[nodeName].
The combination of Pass 1 (with nominated pods) and Pass 2 (without nominated pods) ensures accurate capacity accounting across all phases of the settling window:
- Surplus victim capacity is never treated as free unless
RemovePodremoved the victim pod or t3 has passed: If a victim claimClaim-Vholds more capacity thanPconsumes (for example, two devicesGPU-0andGPU-1whenPonly needsGPU-0), subtractingClaim-Vand addingP’sAllocationResultin Pass 1 leavesGPU-1free in Pass 1, butClaim-V(GPU-0andGPU-1) remains allocated in Pass 2 unlessRemovePod(V)explicitly removedV. Thus, during normalFilter(t0 to t3) or during a preemption simulation between t1 and t3 (whenVis already gone fromNodeInfo.Pods), Pass 2 rejects any attempt to allocateGPU-1whileClaim-Vis still allocated in the API server. - Once
Claim-Vis deallocated at t3:Claim-Vdrops out of the state built byPreFilter, step 1 inAddPod(P)becomes a no-op, andPcontinues to hold onlyGPU-0via step 2 until t4, whileGPU-1becomes immediately allocatable in normalFilter(passing both Pass 1 and Pass 2).
The capacity is held only against pods whose priority is not higher than the preemptor’s. When a pod
of higher priority evaluates the node, RunFilterPluginsWithNominatedPods does not call AddPod
for the lower-priority preemptor, so the higher-priority pod sees the capacity as free once t3
passes and may be allocated it, after which the preemptor must preempt again.
Interaction with the preemption simulation
Two cases illustrate how AddPod composes with subsequent preemption simulations on the same node
while an earlier preemption on behalf of P1 is settling:
Sharing a terminating victim (t0 to t1): Suppose victim
Von NodeNholds a claimClaim-Vwith two devices (GPU-0andGPU-1), andP1(needing one device) preemptsVand records a nomination forGPU-0(with victim claimClaim-Vtracked on NodeN). WhileVis still terminating inNodeInfo.Pods(t0 to t1), an equal-priority podP2(also needing one device) runs a preemption simulation on NodeN:DefaultPreemptioncallsRemovePod(V), removingClaim-V(GPU-0andGPU-1) from the simulated state.- In Pass 1 of
RunFilterPluginsWithNominatedPods,AddPod(P1)addsP1’s nominated allocation (GPU-0).Filter(P2)seesGPU-0occupied byP1andGPU-1free, and computes a simulated allocation ofGPU-1forP2. - Pass 2 (without nominated pods) also succeeds because
RemovePod(V)removedClaim-Vfrom the simulated state. BothP1(GPU-0) andP2(GPU-1) therefore selectVas a victim, withClaim-Vtracked on NodeN, and schedule onceClaim-Vis deallocated at t3. (IfP2instead arrives between t1 and t3 afterVhas already been deleted fromNodeInfo.Pods,DefaultPreemptionsees no victim podVon NodeNto evict;P2simply waits untilClaim-Vis deallocated at t3, when theResourceClaiminformer event wakesP2and schedules it ontoGPU-1in normalFilter.)
Preempting a second victim on the same node after the first victim is deleted (t1 to t3): Suppose an 80Gi device on Node
Nis split betweenV1(Claim-V1, 40Gi) andV2(Claim-V2, 40Gi).P1(needing 40Gi) preemptsV1, and by t1V1has terminated and been deleted fromNodeInfo.PodswhileClaim-V1(40Gi) is still allocated in the API server. When an equal-priority podP2(needing 40Gi) runs a preemption simulation on NodeNbetween t1 and t3,DefaultPreemptioncallsRemovePod(V2)(freeingV2’s 40Gi), but does not callRemovePod(V1)becauseV1is no longer inNodeInfo.Pods. In Pass 1 ofRunFilterPluginsWithNominatedPods,AddPod(P1)subtractsClaim-V1(40Gi) before addingP1’s nominated 40Gi, preventingClaim-V1andP1from being double-counted to 80Gi. Both Pass 1 and Pass 2 therefore see 40Gi in use and 40Gi freed byRemovePod(V2), allowingP2to preemptV2.
Deferring further preemption
While a nomination is live on a node and any tracked victim ResourceClaim on that node still has
Status.Allocation != nil in the informer cache, the preemptor must not start a new preemption.
Together with the existing terminating-pod check, which covers t0 to t1, this covers the settling
window up to t3. Once all tracked victim claims on the nominated node have been deallocated
(Status.Allocation == nil), the preemptor becomes eligible to preempt again; if its placement no
longer works at that point (for example because a higher-priority pod took the freed device),
preempting again is the correct behavior.
PodEligibleToPreemptOthers and prepareCandidate belong to the DefaultPreemption plugin, and we
do not want to make that plugin aware of DRA. We therefore propose a small and generic addition to
the scheduling framework: an optional NominationExtensions interface that manages the nomination
lifecycle and preemption eligibility check (provisional for Alpha while we evaluate moving DRA state
and nominations into the core framework for Beta):
// NominationExtensions is an optional interface for plugins that manage
// resources requiring explicit reservation for nominated pods and/or
// asynchronous reclamation when victim pods are preempted.
type NominationExtensions interface {
Plugin
// AddNominatedPod is called when preemption selects a winning candidate for a pod.
// state is the scheduling cycle's CycleState.
AddNominatedPod(ctx context.Context, state CycleState, pod *v1.Pod, nodeName string, victims []*v1.Pod)
// RemoveNominatedPod is called when a pod's nomination is cleared or superseded.
RemoveNominatedPod(ctx context.Context, pod *v1.Pod)
// PodEligibleToPreempt reports whether a pod with an existing nominatedNodeName
// is eligible to start a new preemption, or whether resources freed by an earlier
// preemption on nodeName are still being reclaimed.
PodEligibleToPreempt(ctx context.Context, pod *v1.Pod, nodeName string) (bool, string)
}
The framework invokes registered plugins implementing NominationExtensions:
AddNominatedPodis called byDefaultPreemptionwhen recording the winning preemption candidate for a pod.RemoveNominatedPodis called by the framework’sPodNominatorwhenever a pod’s nomination is cleared or replaced.PodEligibleToPreemptis called byDefaultPreemption.PodEligibleToPreemptOthersalongside its existing terminating-pod check, stopping at the first plugin that reportsfalseand surfacing the returned reason in the pod’s scheduling condition.
In Alpha, the dynamicresources plugin implements PodEligibleToPreempt by checking whether
nodeName still has any victim ResourceClaim waiting for deallocation
(Status.Allocation != nil)—matching DefaultPreemption’s existing node-scoped check for
terminating pods on nominatedNodeName—while a full cross-preemptor “assumed victim” simulation
mechanism for both pods and ResourceClaims (aligned with Workload-Aware Preemption) is deferred
to Beta.
When the resourceclaim controller clears Status.Allocation on a victim claim at t3, the
scheduler’s ResourceClaim informer event handler invokes SchedulingQueue.MoveAllToActiveOrBackoffQueue,
which calls dynamicresources’s QueueingHint (isSchedulableAfterClaimChange). As a secondary
optimization, the QueueingHint withholds a wake-up (QueueSkip) for the preemptor while any other
tracked victim claim on its nominated node still has Status.Allocation != nil, and returns Queue
as soon as the last victim claim on that node is deallocated. This avoids pointless scheduling
attempts while some victim claims are still settling.
Why simulated allocations and victim claims rather than devices
On preemption, draManager records both the preemptor’s simulated AllocationResults (per
preemptor, to hold capacity via AddPod) and the victim ResourceClaim UIDs (per node, to
track ongoing deallocation between t0 and t3), rather than nominating raw device names:
- Why simulated
AllocationResults rather than device names: With consumable capacity and partitionable devices, a bare device name does not express a capacity share or the counter draw of a partition. AnAllocationResult(*resourceapi.AllocationResult) is the exact data structure produced by the DRA allocator duringFilterand consumed bydynamicresourceswhen building its allocated state. Recording the preemptor’s simulatedAllocationResultholds the exact share or partition counters the preemptor needs, without blocking the rest of the device. - Why victim
ResourceClaimUIDs are also recorded: A simulatedAllocationResulthas no completion condition of its own in the API before t4, and checking whether a simulatedAllocationResultconflicts with arbitrary deallocating claims in the informer cache would require complex counter and partition-overlap math acrossResourceSlices (since a preemptor’s partition name may differ from the partitions held by the victims). Recording the UIDs of the victim claims thatRemovePodreleased during the winning simulation gives an O(1), directly observable completion condition (claim.Status.Allocation == nilat t3) and letsAddPodsubtract those exact victim allocations whileStatus.Allocation != nilwith zero conflict math.
Deferred: workload-aware preemption
Workload-aware preemption (KEP-5710
) removes candidate victims by mutating the scheduler snapshot
and then reruns the pod group scheduling algorithm with fresh CycleStates, instead of calling
PreFilterExtensions.RemovePod for each victim. A plugin that learns about removals only through
RemovePod, as specified here, therefore goes on counting the victims’ devices as allocated, and
the preemption frees nothing. The pod group stays pending, which is the behavior without this
feature. No state is corrupted and no device is allocated twice.
Supporting that path needs three additions that we prefer to design separately:
- The plugin has to rebuild its allocated state by reconciling the ResourceClaims it reads against
the pods present in the
NodeInfosthatPreFilterreceives. Doing so safely requires distinguishing a consumer that has been removed from the snapshot from one that the plugin has merely not observed yet, such as a pod that another scheduler has reserved a claim for but not yet bound. - A nomination has to be owned by the preempting pod group rather than by a single pod. One workload-aware preemption produces a nominated placement for every member of the group, all at the same priority, so per-pod nominations would need to be coordinated across the group.
- Workload-aware preemption (as well as multi-node DRA claims) changes the scope of preemption and
nomination from node-local to cluster-global. Currently, the scheduler applies nominations
per-node during
Filter(RunFilterPluginsWithNominatedPodsonly invokesAddPodfor pods whosenominatedNodeNamematches the node being evaluated). Supporting multi-node DRA claims and global workload-aware preemption requires applying nominations globally before filtering starts (for example duringPreFilter).
Test Plan
[x] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.
Prerequisite testing updates
Unit tests
k8s.io/kubernetes/pkg/scheduler/framework/plugins/dynamicresources: 82.4%
Integration tests
Integration tests will be added that covers the normal functionality of preemption with DRA, with specific tests covering all the special scenarios called out above. For the settling window in particular:
- a preemptor does not preempt a second set of victims while the claims freed by its earlier preemption are still allocated;
- devices freed by a preemption are not allocated to another pod of equal or lower priority before the preemptor has been scheduled;
- a pod of higher priority can be allocated those devices, and the preemptor then preempts again;
- a nomination is removed and held capacity is released when the preemptor is deleted or its
nominatedNodeNameis cleared.
e2e tests
E2e tests will be added to cover the basic scenarios for preemption with DRA. The more complicated scenarios will be handled by integration tests.
Graduation Criteria
Alpha
- Feature implemented behind a feature flag
- Unit, integration and e2e tests completed and enabled
- The cost of the preemption simulation measured with scheduler_perf
Beta
- Tests are in Testgrid and linked in the KEP
- Alignment on whether DRA state and claim nominations should move from
dynamicresourcesinto the core scheduler framework (replacingNominationExtensions) - An agreed-upon design for how claim nominations can be persisted in the API to survive a scheduler restart and be visible to external components such as Cluster Autoscaler.
- Metrics for claim nominations exposed and documented
GA
- 2 examples of real-world usage
- Allowing time for feedback
Upgrade / Downgrade Strategy
Standard upgrade/downgrade strategies may be used, no special configuration changes are needed. The changes are local to the scheduler.
Claim nominations are held only in the scheduler’s memory, so a downgrade discards them without leaving anything behind. Preemptors that were waiting revert to the behavior they would have with the feature disabled.
Version Skew Strategy
This feature is implemented entirely in the scheduler and requires no new behavior from any other component. It does rely on the resourceclaim controller deallocating the ResourceClaims of deleted pods, but that behavior predates this KEP, so there is no version skew concern between the scheduler and kube-controller-manager.
Production Readiness Review Questionnaire
Feature Enablement and Rollback
How can this feature be enabled / disabled in a live cluster?
- Feature gate (also fill in values in
kep.yaml)- Feature gate name: DRAPreemption
- Components depending on the feature gate:
- kube-scheduler
The gate is disabled by default.
Does enabling the feature change any default behavior?
Yes, it will cause the scheduler to preempt pods referencing ResourceClaims to make room for higher priority pods. There is no API change for this feature, so it will potentially impact any Pod in the cluster when the feature is enabled.
Because the gate is disabled by default, upgrading to a release that contains this feature does not by itself change any behavior. An operator has to enable the gate.
Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?
Yes, disabling the feature will prevent the scheduler from preempting pods that reference ResourceClaims.
Disabling it also discards any claim nominations that are currently held by the scheduler in its memory, which releases the capacity being held for preemptors that have not yet been scheduled. Nothing is persisted, so there is no state to clean up and no reconciliation is required.
What happens if we reenable the feature if it was previously rolled back?
That it was previously rolled back have no impact. Reenabling it just means that preemption of pods referencing ResourceClaims will again be considered.
Are there any tests for feature enablement/disablement?
Since this is a purely in-memory feature controlled by a feature gate (which requires a scheduler restart to change), testing with the feature gate enabled and disabled in unit and integration tests is sufficient; no separate enablement/disablement transition tests are needed.
Rollout, Upgrade and Rollback Planning
How can a rollout or rollback fail? Can it impact already running workloads?
A rollout will immediately make pods referencing ResourceClaims eligible for preemption. So if there are higher priority pods that can’t be scheduled due to pending ResourceClaims, the result will be that lower priority pods will get preempted.
Similarly, a rollback means that pods referencing ResourceClaims will no longer be eligible for preemption. But already preempted workload will remain in the pending state until sufficient resources are freed up.
What specific metrics should inform a rollback?
The specific signal that would suggest this feature should be rolled back, would be if pods are being preempted when they shouldn’t be. This means there is a bug somewhere in the implementation.
If scheduler_preemption_victims or scheduler_preemption_attempts_total increases significantly when the feature is enabled, but we don’t
see a corresponding increase in pods being scheduled, we should investigate. It would suggest
that pods are being preempted incorrectly and the higher-priority pods are not actually
being scheduled. A significant increase in scheduler_plugin_execution_duration_seconds for
DefaultPreemption (PostFilter) or DynamicResources (Filter) would also indicate that DRA
preemption simulations are causing a scheduling latency regression.
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
It will be tested by bringing up a KinD cluster and enabling the feature gate. We will then run a simple workload with higher priority pods and then enable the feature gate to verify that the lower priority pods are preempted. We will then disable the feature gate and verify that another pod with an even higher priority does not preempt the existing pod.
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?
No
Monitoring Requirements
How can an operator determine if the feature is in use by workloads?
It is enabled on a cluster-level, so if it is enabled, it is in use. It will only impact pods that are referencing ResourceClaims.
How can someone using this feature know that it is working for their instance?
When a Pod is preempted, it can be observed in the following ways:
- Events
- Event Reason: Preempted
- API .status
- Condition name: DisruptionTarget
Seeing this on a Pod that references ResourceClaims shows that the feature is working.
What are the reasonable SLOs (Service Level Objectives) for the enhancement?
Kubernetes does not have SLOs for preemption, so neither will this enhancement.
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
- Metrics
- Metric name:
scheduler_preemption_victimsandscheduler_preemption_attempts_total. These are not specific to DRA and carry no DRA-specific labels, so they can be compared across the cluster before and after the gate is enabled. - Metric name:
scheduler_plugin_execution_duration_seconds(withplugin="DefaultPreemption",extension_point="PostFilter"andplugin="DynamicResources",extension_point="Filter") to monitor the latency cost of DRA preemption simulations. - Components exposing the metric: kube-scheduler
- Metric name:
Are there any missing metrics that would be useful to have to improve observability of this feature?
Dedicated metrics for tracking active DRA claim nominations and how they resolve (e.g. scheduled vs.
cleared) will be added for Beta once we have finalized whether claim nominations remain in
dynamicresources or move into the core framework.
Dependencies
Does this feature depend on any specific services running in the cluster?
Yes. The resourceclaim controller in kube-controller-manager deallocates the ResourceClaims of preempted pods, and preemption for DRA cannot complete until it has done so. If the controller is unavailable or lagging, preemptors remain unschedulable and wait for their nominated claims to be deallocated; see What are other known failure modes? .
Scalability
Will enabling / using this feature result in any new API calls?
The dynamicresources plugin itself makes no new API calls. The preemption simulation runs against
state the plugin already maintains.
Enabling the feature does however make pods referencing ResourceClaims eligible for preemption, so
the scheduler will issue pod deletions and DisruptionTarget condition updates for those pods
where previously it issued none. The rate is proportional to the rate of preemption, and the cost
per preemption is the same as for pods that do not use DRA.
Will enabling / using this feature result in introducing new API types?
No
Will enabling / using this feature result in any new calls to the cloud provider?
No
Will enabling / using this feature result in increasing size or count of the existing API objects?
No
Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?
No. It might lead to increase time taken to check if lower priority pods can be preempted to make room for a higher priority pod, since we don’t do that today.
Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?
It can lead to some additional work in the scheduler, since we are enabling preemption simulation for a new scheduler plugin.
The scheduler additionally holds one claim nomination per preemptor pod that is waiting for an
earlier preemption to settle, containing the simulated allocation results for the preemptor and the
UIDs of the victim claims being released. The number of such nominations is bounded by the number of
in-flight preemptions, and each is discarded when the preemptor is scheduled, deleted, or has its
nominatedNodeName cleared.
Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?
No
Troubleshooting
How does this feature react if the API server and/or etcd is unavailable?
The dynamicresources plugin does not add any calls to the API server compared to what it already
does, so the plugin behaves as it would without this feature enabled.
Preemption as a whole cannot make progress while the API server is unavailable: victim pods cannot be deleted and their ResourceClaims cannot be deallocated, so claim nominations remain active until the API server recovers and the victim claims are deallocated.
What are other known failure modes?
- The resourceclaim controller does not deallocate a nominated claim, for example because it is
unhealthy. The preemptor remains unschedulable and its nomination continues to hold the simulated
capacity until the controller recovers and deallocates the claim (or the preemptor is deleted).
This is visible in the API as the preemptor pod remaining Pending with
.status.nominatedNodeNameset while the victim ResourceClaims retainstatus.allocationafter their pods have been deleted; operator attention is required to restore the controller. - The scheduler restarts while preemptions are settling. Nominations are in memory and are lost, so the affected preemptors may have their devices taken by another pod or may preempt again. The effect is limited to preemptions that were in flight at the time of the restart.
What steps should be taken if SLOs are not being met to determine the problem?
There are no SLOs for this feature.
Implementation History
- 1.37: first revision of the KEP and an initial implementation. Deferred to 1.38 over the handling of asynchronous device reclamation and the interaction with workload-aware preemption.
- 1.38: revised with the claim nomination and preemption settling designs, which address the asynchronous reclamation problem, and scoped to the pod-by-pod preemption path. Support for workload-aware preemption is deferred to a later revision. Targeted at Alpha.
Drawbacks
It complicates the logic in the dynamicresources plugin and can lead to slower preemption when pods are using DRA.
Holding capacity for a nominated preemptor can also leave devices idle for the duration of the settling window if the preemptor is deleted or fails before being scheduled. See Risks and Mitigations .
Alternatives
Kubernetes has a well-established framework for preemption, and this KEP uses it rather than adding a separate mechanism for DRA. The alternatives below all concern the settling window described in Preemption settling for DRA .
Nominating devices rather than simulated allocations and victim claims. Discussed in Why simulated allocations and victim claims rather than devices .
Holding the victims’ claims rather than the preemptor’s simulated allocation. Rather than
recording the simulated AllocationResults computed for the preemptor and adding them via AddPod,
the plugin could hold the AllocationResults of the victim claims that were released. We rejected
this because a victim claim may hold more capacity than the preemptor needs—such as multiple devices
when the preemptor needs only one, or a larger share of consumable capacity—which would unnecessarily
block other pods from claiming the surplus capacity during the settling window.
Allocating the preemptor’s claims eagerly. Rather than recording an in-memory nomination, the scheduler could write the allocation computed during preemption to the preemptor’s ResourceClaims immediately and mark it as not yet usable. We rejected this because the victims’ claims still hold the same devices until t3, so the two allocations would conflict; because the allocation was computed against a simulated state and may no longer be valid once the cluster state converges; and because it would need an unwind path for the case where the preemptor ultimately cannot be scheduled.
Deallocating the victims’ claims from the scheduler. The scheduler could clear
Status.Allocation itself once the victim pods are gone, instead of waiting for the resourceclaim
controller. This shortens the interval from t1 to t3 but does not remove it, does nothing for the
grace period from t0 to t1, and does not prevent device stealing between t3 and t4. It would also
have to cope with the scheduler’s informer cache not yet reflecting its own update. It is a
possible later optimization rather than a substitute.
Waiting for the victims to terminate within the scheduling cycle. This removes the settling window altogether, but blocks the scheduler for the length of the victims’ grace periods and is therefore not acceptable.
Deferring preemption from the plugin’s PostFilter rather than through NominationExtensions.
The dynamicresources plugin is registered immediately before DefaultPreemption at the PostFilter
extension point, and RunPostFilterPlugins returns as soon as a plugin reports
UnschedulableAndUnresolvable. The plugin could therefore return that status while the pod has an
unsettled nomination, and DefaultPreemption would not run for that scheduling attempt. We did not
choose it because it still would not notify the plugin which candidate won SelectCandidate, and
because it depends on the relative ordering of the two plugins at PostFilter.
Accepting the behavior and relying on priority. Preemption is best-effort, so the scheduler could simply let the preemptor retry. We rejected this because the situation that motivates this KEP, many pods contending for scarce accelerators, is precisely the one in which a preemptor can be starved indefinitely while each of its attempts evicts a further set of victims.
Infrastructure Needed (Optional)
No