KEP-5981: DRA Sharing Affinity
KEP-5981: DRA Sharing Affinity
- Release Signoff Checklist
- Summary
- Motivation
- Proposal
- Design Details
- Production Readiness Review Questionnaire
- Implementation History
- Drawbacks
- Alternatives
- CEL selector on each extractor
- Per-device extractor reference
- Well-known JSON schema inside OpaqueDeviceConfiguration
- Typed
Structuredsibling onDeviceConfiguration - Claim-side-only SharingAffinity (on DeviceRequest)
- Object Reference-based Affinity Matching
- CEL on runtime scheduler lock state (rejected variant)
- Future Enhancements
- Infrastructure Needed
Release Signoff Checklist
Items marked with (R) are required prior to targeting to a milestone / release.
- (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
- (R) KEP approvers have approved the KEP status as
implementable - (R) Design details are appropriately documented
- (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
- e2e Tests for all Beta API Operations (endpoints)
- (R) Ensure GA e2e tests meet requirements for Conformance Tests
- (R) Minimum Two Week Window for GA e2e tests to prove flake free
- (R) Graduation criteria is in place
- (R) all GA Endpoints must be hit by Conformance Tests
- (R) Production readiness review completed
- (R) Production readiness review approved
- “Implementation History” section is up-to-date for milestone
- User-facing documentation has been created in kubernetes/website , for publication to kubernetes.io
- Supporting documentation—e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes
Summary
This KEP proposes an extension to Dynamic Resource Allocation (DRA) that allows
the kube-scheduler to handle resources that are conditionally fungible.
KEP-5075 (Consumable Capacity)
introduced the ability to track numerical capacity (e.g., 16 slots of a NIC)
and share devices across multiple claims via allowMultipleAllocations.
However, it assumes all claims are fungible—any claim can share the device with
any other claim.
Real-world hardware is often modal (i.e., once partially allocated, it must operate in a single configuration mode for all of its current consumers): the device requires all subsequent consumers to share a specific configuration. For example:
- Multi-pod NIC sharing: A network DRA driver shares a NIC across 16 pods, but all pods must belong to the same subnet. Once the first pod configures the NIC for Subnet A, the remaining 15 slots are restricted to Subnet A.
- FPGA bitstream sharing: An FPGA can serve multiple inference pods, but all must use the same bitstream. Once bitstream-ml-v2 is loaded, other pods needing bitstream-crypto-v1 must use a different FPGA.
This KEP introduces a sharingAffinity field on ResourceSlice.spec
that allows drivers to declare, via CEL expressions, how to extract
the affinity keys that constrain sharing from the driver’s existing
opaque device configuration. sharingAffinity is published as
pool-level metadata in dedicated slices within the pool — a
slice that carries sharingAffinity does not carry devices, and
vice versa — mirroring the established sharedCounters (KEP-4815)
pattern. A pool may publish one or more such slices; the scheduler
takes the union of their extractors, so the schema declaration stays
scoped to the pool even when the pool is chunked across many
ResourceSlices.
On the claim side, there is no API change — workloads continue to
author their driver’s opaque config exactly as they do today. When
the scheduler is evaluating a candidate device, it looks up the
pool’s metadata slice, runs every published CEL extractor
against the opaque-config objects in scope for the request
currently being filtered (per DeviceClaimConfiguration.Requests),
and — if every extractor is satisfied —
records the resulting key/value pairs in AllocatedState
alongside consumed capacity. This enables the scheduler to gate remaining capacity
on locked devices and safely reuse them for compatible claims when
selected by the existing allocator. Alpha provides correctness only —
affinity-aware preference (packing) is delivered in beta; see
Goals
.
In addition, if a device already has active allocations whose
affinity cannot be reconstructed (for example, legacy claims created before the
feature was enabled), the scheduler treats that device conservatively and does
not place new sharingAffinity allocations on it until the device becomes
clean.
sharingAffinity in this KEP refers specifically to compatibility for
co-allocation on a shared device; it is distinct from pod affinity,
anti-affinity, or topology-aware placement.
Motivation
As AI and HPC workloads move toward higher density, hardware partitioning (SR-IOV, GPU slicing, FPGA multi-tenancy) is becoming standard. These physical devices often have a “modal” constraint (see Summary for the definition and concrete examples).
Currently, the scheduler is unaware of this “lock.” It may schedule a Pod requiring a different configuration to the same device because it sees “available capacity.” In short: In these scenarios, Quantitative Sharing (how many slots?) fails without Qualitative Gating (what mode are those slots in?). This leads to:
- Allocation failures at the node level: The driver rejects incompatible binds at prepare time, after the scheduler has already committed
- High scheduling latency: The scheduler retries the same failing combination, thrashing between candidates
- Resource starvation: Without affinity awareness, same-subnet pods spread across multiple devices instead of consolidating—wasting capacity
- Complex driver workarounds: Drivers resort to placeholder patterns with race conditions and ResourceSlice churn (see Status Quo below)
The scheduler’s AllocatedState currently tracks consumed capacity but not the
affinity values that determine sharing compatibility. This KEP closes that gap.
Status Quo: Driver-Side Placeholder Pattern
Without this KEP, drivers must use a “placeholder pattern” today:
- Publish devices with
capacity: 1initially - Wait for first claim to determine affinity value
- Update ResourceSlice with actual capacity, writing the affinity value into the device’s
attributesmap - Use CEL selector to match against that
attributesentry
Problems:
- Race condition: Second pod may go to different device before expansion
- ResourceSlice churn: Constant updates as pods come and go
- Driver complexity: State machine for expand/contract lifecycle
Goals
- Enable the scheduler to gate remaining capacity on a device based on a required affinity key
- Provide a mechanism for drivers to signal compatibility requirements for
shared hardware via
sharingAffinityon the ResourceSlice — without requiring drivers to change their existing opaque config schemas, and without requiring workload authors to learn a new claim-side API - Reduce fragmentation of cluster resources by enabling the scheduler to
pack workloads with compatible sharing requirements onto already-locked
devices (delivered in beta as a sharing-affinity term added to
DynamicResources.computeScore— see Affinity-aware scoring (planned for Beta) ; alpha provides correctness only) - Track affinity values in
AllocatedStateso subsequent scheduling decisions respect the first claim’s lock-in - Maintain backward compatibility with devices that have no sharing affinity constraints, and with existing opaque-config workloads
Non-Goals
- Defining hardware-specific affinity key names (these remain driver-defined)
- Managing the physical lifecycle of the device configuration (this remains the driver’s responsibility)
- Changing how capacity is tracked (that’s KEP-5075)
- Supporting affinity across multiple devices. The lock is scoped to a
single
Deviceobject inResourceSlice.spec.devices[]— affinity is never shared across separateDeviceobjects, even ones of the same type within the same pool, and certainly not across device types or pools. - Retrofitting affinity-aware sharing onto already-in-use devices when active claims do not expose reconstructable affinity values. In alpha, such devices are treated conservatively until they drain clean.
- Guaranteeing lock-breaking preemption. This KEP does not introduce any preemption logic. Baseline DRA preemption is being introduced separately by KEP-5690 ; this KEP is designed to compose transparently with it (see Composition with DRA Preemption (KEP-5690) ) but does not depend on or deliver it.
Proposal
Add a sharingAffinity field to ResourceSlice.spec that publishes
a list of driver-defined CEL extractors. Each entry names one
affinity key and carries the CEL expression that produces its value,
and every entry applies to every device in the pool. The scheduler
evaluates an entry’s expression once for each opaque config in the
claim, binding that config to a CEL variable named object — its JSON
payload decoded into a nested map. (DRA already requires the payload
to be valid JSON, so this KEP adds no new constraint.) The expression
returns either a string value — the affinity key for that config — or
the empty string "", meaning the extractor does not apply to it. A
request satisfies an extractor when it produces a non-empty value
across the request’s in-scope opaque configs, and the device is viable
only if every extractor in the pool is satisfied. The non-empty
returns from all extractors form the request’s effective
affinity key map. If two opaque-config objects in the same request
return different non-empty values for the same key, the claim is
self-inconsistent and is treated as a user authoring error (see
Extraction failure modes
).
Sharing affinity is published in dedicated metadata slices
that are mutually exclusive with devices. Device-bearing slices in
the same pool reference back to the metadata via the standard pool
tuple (driver, pool.name, pool.generation). A pool may
publish one or more metadata slices; the scheduler takes the union
of sharingAffinity extractor entries across all metadata slices in
the same complete pool (same generation, all resourceSliceCount
members present, per the pool-completeness rule from
KEP-4815
).
Before reading metadata, the scheduler invokes the existing
pool-completeness check; devices in incomplete or invalid pools are
treated as ineligible for sharing-affinity decisions until the pool
becomes complete and clean. Conflict resolution rides on existing
rules — incomplete or invalid pools are skipped entirely
(KEP-4815’s fail-closed gate), and key names must be unique across
all of the pool’s extractors (see
Key name uniqueness
), so
multi-metadata-slice publishing reduces to the same semantics as a
single metadata slice.
# Metadata slice: carries sharingAffinity, no devices.
apiVersion: resource.k8s.io/v1
kind: ResourceSlice
metadata:
name: networking-node-a-meta
spec:
driver: networking.example.com
nodeName: node-a
pool:
name: node-a
generation: 7
resourceSliceCount: 2
sharingAffinity:
- name: subnet
expression: |
object.apiVersion == "networking.example.com/v1"
&& object.kind == "NICConfig"
? object.subnetID : ""
- name: pkey
expression: |
object.apiVersion == "networking.example.com/v1"
&& object.kind == "NICConfig"
? object.ibPKey : ""
---
# Device slice: carries devices, no sharingAffinity. Joined to the
# metadata slice by the (driver, pool.name, generation) tuple.
apiVersion: resource.k8s.io/v1
kind: ResourceSlice
metadata:
name: networking-node-a-0
spec:
driver: networking.example.com
nodeName: node-a
pool:
name: node-a
generation: 7
resourceSliceCount: 2
devices:
- name: eth1
allowMultipleAllocations: true
capacity:
networking.example.com/slots:
value: "16"
Each CEL expression is responsible for returning the empty string when
the opaque-config object is not one it applies to. How it decides is
the driver’s choice; guarding on object.apiVersion and object.kind
is the recommended convention, not a requirement. This contract
gives drivers full control over which opaque-config schemas a given
extractor applies to, and avoids forcing the scheduler to know
anything about driver schema taxonomies. CEL runtime errors (for
example, dereferencing a field that does not exist on the parsed
config) are not equivalent to returning "" — they are ResourceSlice
authoring errors, and the scheduler aborts allocation and fails
scheduling for the Pod rather than skipping the device (see
Extraction failure modes
). An
apiVersion/kind guard is the simplest way to avoid this for configs
that belong to other schemas.
If a driver supports multiple opaque-config schemas where the extracted fields live at the same path on each, one extractor can handle them all:
sharingAffinity:
- name: subnet
expression: |
(object.apiVersion == "networking.example.com/v1" ||
object.apiVersion == "networking.example.com/v1beta1")
&& object.kind == "NICConfig"
? object.subnetID : ""
All extractors published by a pool apply to every device in that pool; there is no per-device dispatch. A pool is therefore homogeneous with respect to sharing affinity — every device in it is gated on the same set of affinity keys. A driver whose devices need different affinity keys publishes them in separate pools.
A claim continues to author the driver’s opaque config exactly as it does today — no new claim-side API field is introduced:
config:
- requests: ["nic"]
opaque:
driver: networking.example.com
parameters:
apiVersion: networking.example.com/v1
kind: NICConfig
vendor: acme
subnetID: subnet-A
ibPKey: "0x8001"
qos: gold # driver-private; not extracted, not seen by scheduler
mtu: 9000 # driver-private
vlanTag: 100 # driver-private
When the scheduler evaluates a multi-allocatable device backed by a
pool that declares sharingAffinity:
- First request targeting this device: When evaluating a
candidate device in this pool for a
DeviceRequest, the scheduler looks up the pool’s metadata slice and, for eachsharingAffinityentry, runs the entry’s CEL expression against every opaque config object that is in scope for the request (perDeviceClaimConfiguration.Requests). An extractor is satisfied when it produces a non-empty return across the request’s opaque configs. The device is a viable option for the request only if every extractor is satisfied; the union of their key/value contributions is the request’s effective affinity map and is recorded inAllocatedStatealongside consumed capacity. - Subsequent requests: The scheduler runs the same
extraction for each new request targeting devices in the pool
and compares the merged key set to
AllocatedState’sLockedAffinity. - Mismatch: If extracted keys do not match the recorded keys — as defined under Extraction failure modes — the device is not a viable option for that request and the scheduler tries another candidate device. Key sets always align, because Strict Gating admits a request only when every extractor is satisfied: a valid config must set every field the pool extracts.
- Match: If all extracted keys match and capacity is available, allocation proceeds.
Keys absent from the slice’s CEL expression set are not extracted, so
driver-private config (e.g., qos, mtu, vlanTag above) flows
through to NodePrepareResources untouched and never participates in
lock evaluation. The driver remains the sole authority for those
fields.
Alpha Design Decisions
1. Placement of extraction: ResourceSlice (driver-side, pool-level metadata slice)
sharingAffinity lives on ResourceSlice.spec and is mutually
exclusive with devices. A pool publishes one or more metadata
slices that carry sharingAffinity and zero devices; the pool’s
device-bearing slices carry devices and no sharingAffinity. When
multiple metadata slices exist in the same complete pool, the
scheduler unions their extractor entries (per the multi-slice rule
described above).
The driver is the natural owner: the extraction logic describes the driver’s own opaque config schema, which is uniform across all devices the driver publishes in this pool and evolves with driver versions, not with hardware instances. Pool-level placement avoids per-slice duplication, removes the drift risk of having to keep the extractor block in sync across many slices when a pool is chunked, and aligns the declaration site with what is being declared (a driver-schema property, not a device property). Per-device overrides can be added later if a driver ever needs them.
2. How affinity values reach the scheduler: CEL on opaque config
The scheduler does not interpret driver-private fields — it only
evaluates the CEL expressions the driver has explicitly published as
extraction logic. Each extractor is responsible for recognizing the
schema(s) it applies to (typically by inspecting object.apiVersion
/ object.kind and returning "" when it does not).
// SharingAffinityExtractor declares one CEL expression that produces
// one sharing-affinity key from a driver's opaque device config.
// Every entry in a pool's `sharingAffinity` list applies to every
// device in the pool. See Design Details → API Enhancement for the
// canonical godoc.
type SharingAffinityExtractor struct {
Name string // affinity key name; unique within the pool
Expression string // CEL over `object`; must return a string
}
Properties of this design:
- No claim-side API change: workloads keep authoring the driver’s existing opaque config. No migration cost on the user side.
- Minimal scheduler assumption: the scheduler’s only assumption is “run some CEL against the claim’s opaque configs and see what key/value pairs come back.” It never owns or interprets the driver’s schema.
- Driver-private keys are invisible to the scheduler: anything the driver does not publish a CEL expression for is never extracted. The scheduler cannot inadvertently lock on hardware-configuration fields (qos, mtu, vlan) that drivers want to keep within their domain.
Extractor evaluation contract
An extractor’s CEL is evaluated against every opaque-config object in scope for the request being filtered. For each (extractor, opaque-config, key) triple:
- A non-empty string return contributes that
key → valueto the request’s effective affinity map. - An empty string return contributes nothing for that key from that object. This is the idiom drivers use to express “this extractor does not handle this opaque-config schema.” Empty returns are not errors.
After evaluating every extractor against every opaque-config object, the scheduler checks whether each extractor is satisfied — it produces a non-empty return across the request’s opaque configs. The device is a viable option for the request only if every extractor in the pool is satisfied (Strict Gating).
Extraction failure modes
Three categories of outcomes can arise when evaluating a candidate
device whose pool declares sharingAffinity for a request.
They differ in who is responsible for the failure and what
action the scheduler takes:
Strict Gating outcome (normal, not an authoring error) — at least one extractor is not satisfied for this (request, candidate device) pair (the extreme case is “no keys produced at all”; a partial case is a pool whose
subnetextractor resolved but whosepkeyextractor did not). The device is not a viable option for the request — the safe default, since the driver declared that sharing requires scheduler-readable extraction across every published extractor. The request stays Pending; the pod may still schedule onto another candidate device in this or another pool, including pools that do not declaresharingAffinity.ResourceClaim authoring errors →
UnschedulableAndUnresolvable— the claim’s own opaque configs are internally inconsistent; no slice the driver might publish can fix this.- Self-inconsistent extraction for a request: two
(extractor, opaque-config) pairs in scope for the request
produce the same key with different non-empty values. The
pod is marked
UnschedulableAndUnresolvable; the failure scope is the ResourceClaim, not a single device, because every device backed by asharingAffinitypool that runs the same extractor against the same in-scope opaque configs will see the same contradiction. An Event is emitted on the pod for diagnosability. The user must fix the claim’s opaque configs before rescheduling.
- Self-inconsistent extraction for a request: two
(extractor, opaque-config) pairs in scope for the request
produce the same key with different non-empty values. The
pod is marked
ResourceSlice authoring errors → abort allocation, fail scheduling for the Pod — the slice’s CEL is the problem; the driver must republish the slice to fix it.
- Duplicate key name across extractors: two extractors in the pool declare the same key name. The lock state would be ambiguous about which extractor’s value to store. Within a single slice this is rejected at admission; across metadata slices in the same pool it is caught by pool validation. See Key name uniqueness .
- CEL evaluation error or non-string return: a CEL expression fails to evaluate or returns a non-string. The scheduler does not silently lock the device to an empty key set.
- Cost-budget exhaustion: CEL evaluation exceeds the per-eval cost budget.
- Missing field on the parsed object: a CEL expression
dereferences a field absent from
object(e.g.,object.subneton an object with nosubnet), raising a CEL runtime error. Indicates the extractor’s CEL did not defensively guard the access (e.g., should behas(object.subnet) ? object.subnet : ""). Typical causes: typo (wrong field name) or schema evolution where the driver renamed or removed a field across versions. Inapplicable-object cases (driver-tuning config, config for a different driver kind) should be expressed viahas()guards returning"", not surfaced as errors.
In every ResourceSlice-author-error case the scheduler aborts allocation and fails scheduling for the Pod instead of skipping the device and trying the next candidate. Skipping would emit an Event per candidate device and degrade silently into “no device matched”. The trade-off is that one malformed expression blocks the whole pool; admission-time validation of CEL syntax, cost, and length catches most errors before a slice is accepted.
The canonical extractor idiom — guard with
apiVersion/kind(orhas(...)) and return""from the non-applicable branch — distinguishes “this object intentionally does not contribute” ("") from “this object should have contributed but the schema does not match” (runtime error). Only the latter is treated as a slice authoring error.
Alpha scope
Alpha fully resolves the design around driver-side slice-level
extraction described above. Claims do not control lock-setting
behavior: any compatible claim may establish the initial lock on a
clean device. Claim-side lock-setting policy (for example,
CanSetLock/NeverSetLock) is deferred to Future
Enhancements
.
Alpha standardizes driver-declared CEL extraction, a single feature
gate that governs both the slice-level field and the scheduler logic,
and correct lock enforcement on already-locked devices — but
intentionally stops short of affinity-aware scoring (planned as a
beta-scope contribution to DynamicResources.computeScore, see
Affinity-aware scoring (planned for
Beta)
).
Alpha limitations
Alpha enforces lock compatibility but does not preempt incompatible lock-holders. A higher-priority Pod requiring a different affinity value than the current lock will remain unschedulable on that device until the lock-holder exits or a compatible alternative appears. This is consistent with baseline DRA, which has no preemption support in the absence of KEP-5690 . See Composition with DRA Preemption (KEP-5690) for how this KEP is structured to benefit transparently when KEP-5690 is present.
Alpha also restricts CEL return values to string. List-valued
return types (a claim accepting any of several values for one key)
are a recognized future extension but are out of scope for alpha —
all motivating use cases (subnet IDs, PKeys, bitstream identifiers,
NUMA tags) are single-valued.
A related extensibility note (raised by @pohly in PR review): a
scalar string return type forecloses carrying any metadata
alongside the extracted value. A future revision could return a
structured object (for example a CEL message/struct with fields
such as value, qualifiers, or per-key options) so that the
contract can grow without another API break. Alpha deliberately
stays on string because every motivating use case is a single
opaque identifier and a struct return adds CEL-side type plumbing
and scheduler-side parsing that is not justified by current
requirements; the option is preserved by treating the return type
as a versioned part of the extractor contract rather than a free
parameter. See Future Enhancements
for the
extension path.
User Stories
Story 1: RDMA Partition Key Alignment
A user runs a distributed training job where every Pod must share the same
RDMA Partition Key (PKey) to communicate. The NIC supports 16 VFs. The driver
publishes sharingAffinity on the slice with a CEL expression pulling the
PKey out of its existing opaque config (e.g., pkey: "object.ibPKey"). The
scheduler only co-allocates Pods whose claimed PKey matches the NIC’s current
lock (or selects an unlocked NIC and establishes the lock from the first claim).
- Pod A (pkey-0x8001) is allocated to mlx5_0 → mlx5_0 is now locked to pkey-0x8001
- Pod B (pkey-0x8001) arrives → matches affinity, is eligible to share mlx5_0
- Pod C (pkey-0x8002) arrives → affinity mismatch on mlx5_0; mlx5_0 is not a viable option for Pod C; Pod C is allocated to mlx5_1 instead
Story 2: FPGA Bitstream Sharing
An inference service uses FPGAs to accelerate a specific model. Loading a
bitstream takes several seconds. The driver publishes a CEL expression on
the slice that extracts the bitstream identifier from its opaque config
(e.g., bitstream: "object.bitstreamID"). The scheduler only co-allocates
Pods that request a compatible bitstream onto an already-locked FPGA;
affinity-aware preference for FPGAs that already have the bitstream loaded
(over fresh ones) is delivered in beta.
- Pod A (bitstream-ml-v2) is allocated an FPGA → FPGA locks to bitstream-ml-v2
- Pod B (bitstream-ml-v2) arrives → eligible to share the same FPGA
- Pod C (bitstream-crypto-v1) arrives → the locked FPGA is not a viable option for Pod C; uses a different FPGA or waits
Story 3: Single-subnet NIC Sharing
A network DRA driver advertises NICs that can be shared across up to 16 pods,
but only if pods belong to the same subnet. The driver publishes a CEL
expression on the slice that extracts the subnet from its opaque config
(e.g., subnet: "object.subnetID").
- Pod A (subnet-X) is allocated to eth1 → eth1 is now locked to subnet-X
- Pod B (subnet-X) arrives → matches affinity, is eligible to share eth1
- Pod C (subnet-Y) arrives → affinity mismatch on eth1; eth1 is not a viable option for Pod C; Pod C is allocated to eth2 instead
Notes/Constraints/Caveats
- Affinity is set by the first compatible claim on a clean device: Once a device is allocated with an affinity value, that value is locked until all claims release the device.
- Extractors: The pool’s
sharingAffinitylists one or more extractors, each applying to every device in the pool. Evaluation semantics (Strict Gating across all of the pool’s extractors) are detailed in Filter Phase and Device Selection and Key name uniqueness . - Driver-private keys are invisible to the scheduler: Keys not present
in the slice’s
sharingAffinitylist (e.g., qos, mtu, vlan) are not extracted and do not participate in lock evaluation. They flow through opaquely to the driver atNodePrepareResourcesas today. This is the mechanism by which the scheduler stays out of driver-private config. - String-only matching in alpha: CEL expressions must return
stringvalues. Non-string returns (numbers, booleans, lists, objects) are treated as extraction failures. Workloads use the normalized string form of their identifiers (subnet IDs, PKey hex strings, bitstream names, FQDNs). - CEL evaluation errors: A CEL expression that returns a non-string, errors, or exceeds the per-evaluation cost budget is a ResourceSlice authoring error: the scheduler aborts allocation and fails scheduling for the Pod (see Extraction failure modes ). The scheduler never silently establishes an empty lock when extraction fails.
- Devices opting out of SharingAffinity extraction: Devices in slices without
sharingAffinitybehave as before — any claim can share them regardless of opaque config content. - Legacy allocations with unknown affinity are conservative in alpha: If a device has active allocations for which the scheduler cannot reconstruct the required affinity values (for example, claims created before the feature was enabled, or whose opaque configs produce no non-empty extractor returns under the currently-published CEL), that device is treated as having unknown affinity state and is filtered out for new sharing-affinity scheduling until it becomes fully clean.
Composition with DRA Preemption (KEP-5690)
This KEP does not deliver preemption. Lock-breaking preemption is not a KEP-5981 deliverable in any milestone.
The design is structured so lock-breaking falls out transparently
when KEP-5690 (DRA Preemption)
is present in the cluster: affinity locks live in the same
dynamicresources.stateData structure that KEP-5690 mutates via
AddPod/RemovePod. The only implementation requirement on the
KEP-5981 side is that AddPod/RemovePod correctly mutate
affinity-lock state alongside capacity — which is already required
for KEP-5690 compatibility.
Known limitation: affinity-blind reprieve ordering.
SelectVictimsOnNode’s reprieve order is (priority, PDB, runtime)
and does not consider which Device a Pod’s claim is allocated to.
When a node has multiple candidate devices each holding a lock to a
different value, with asymmetric lock-holder counts, the algorithm
may converge on a victim set on the larger-lock-count device when
sacrificing the smaller one would have sufficed. Always correct,
sometimes over-evicts. A DRA-aware reprieve-ordering hook in
DefaultPreemption would resolve this cleanly but is its own
scheduler-framework enhancement, out of scope for this KEP.
Reclaiming an affinity-locked device is addressed by workload-aware preemption (cluster-wide victim identification), which is out of scope for KEP-5690’s initial integration and left to future enhancements.
If KEP-5690 is not present or is disabled, this KEP provides no preemption capability — see Preemption Cannot Break Affinity Locks .
Handling Legacy Claims with Unreconstructable Affinity
| Device State | New Claim | Result |
|---|---|---|
| 5 legacy claims, affinity unknown | Claim whose CEL extraction yields subnet: A | Filtered out. Existing allocations have unknown affinity, so no new sharing-affinity lock may be established yet. |
| 5 legacy claims, affinity unknown | Claim whose extraction yields no non-empty keys | Filtered out. Missing required scheduler-readable affinity information. |
| Legacy claims drained; device now clean | Claim whose CEL extraction yields subnet: A | Lock set to subnet: A; device now locked. |
Device locked to subnet: A | Claim whose CEL extraction yields subnet: A | Allowed (values match). |
Device locked to subnet: A | Claim whose CEL extraction yields subnet: B | Filtered out (mismatch with lock). |
| All claims released | — | Device fully clean and eligible to establish a new lock. |
Legacy claims continue to run and are not evicted. However, until all unknown allocations on a sharing-affinity device are released, the scheduler does not assume it knows the device’s effective modal state.
Changing sharingAffinity on a Slice with Active Allocations
In alpha, mutating a slice’s sharingAffinity (adding, removing, or
changing extractor entries or CEL expressions) while bound claims
reference devices in that slice is not supported. On the next
reconciliation or scheduler restart, the scheduler re-runs the new CEL
against each bound claim’s opaque config. If any bound claim no longer
yields the same key set as the current LockedAffinity (because a key
was added, removed, renamed, or the CEL now returns a different
value), the device is marked AffinityStates[deviceID].Status = AffinityStatusUnreconstructable
and filtered out for new sharing-affinity scheduling until all such
claims drain.
This is the same conservative-fallback behavior used for legacy claims and
is the deliberate alpha trade-off: the scheduler refuses to silently
downgrade its safety guarantee in the face of an asymmetric extraction
change. Drivers that need to evolve sharingAffinity for in-use slices
should drain affected devices before publishing the change, or stage
schema evolution by adding a new extractor entry (e.g., one whose CEL
guards on a new apiVersion) and migrating workloads to author against
the new schema before retiring the old extractor.
Driver responsibility: drivers should avoid hot-swapping
sharingAffinity entries on slices with active allocations. When
extraction changes are unavoidable (e.g., a hardware capability evolves),
drivers should expect the affected devices to be ineligible for new
affinity-aware scheduling until they drain clean, and should plan rollouts
accordingly (for example, by cordoning the device or rolling out the
extraction change as part of a node reimage).
Compatibility Matrix
To clarify the interaction between requests and devices, the following matrix
outlines how the scheduler and driver evaluate candidates based on whether
the device’s pool declares a sharingAffinity Affinity Extractor list
(AE) and whether the request’s opaque configs
satisfy every extractor in that pool’s
sharingAffinity list (GS — under Strict Gating every extractor
in the pool must be satisfied):
| Scenario | Pool AE | Request GS | Scheduler Outcome | Driver Outcome |
|---|---|---|---|---|
| Standard Feature Use | Yes | Yes | Match enforced. Extracted values match lock + capacity available → request scheduled. | Validates hardware mode matches claim config at NodePrepareResources. Rejects if stale or inconsistent. |
| Strict Gating | Yes | No | Not viable. Device is not a viable option for the request — at least one extractor was not satisfied by the request’s opaque configs (anything from “no keys at all” to “one key missing”). Request stays Pending; may schedule elsewhere. | N/A — request never reaches the driver for this device. |
| Legacy Device Transition | Yes (newly added) | Yes | Not viable while legacy claims are active (Status: Unreconstructable). Allowed once device drains clean. | Validates as normal once request reaches the driver. During transition, driver continues serving legacy claims. |
| Permissive Sharing | No | Yes | Allowed. Pool has no sharingAffinity; opaque config is not evaluated for affinity. Standard capacity matching applies. | Must enforce hardware compatibility independently. Scheduler provides no affinity gating for this device. |
| Legacy/Basic | No | No | Allowed. Standard DRA capacity and attribute matching. | Must enforce hardware compatibility independently. This is the pre-KEP-5981 behavior. |
| Gate Disabled in Scheduler | Yes | N/A | Not viable. The scheduler cannot evaluate affinity, so every device in the pool is skipped rather than allocated with enforcement silently off (see Feature Gates ). | N/A — request never reaches the driver for this device. |
The top rows show the scheduler as the primary enforcer with the driver as a backstop. The bottom rows show the driver as the sole enforcer with the scheduler being permissive. The transition row shows the scheduler being conservative (filtering) while the driver continues serving existing workloads. The last row is the fail-closed case: a pool asked for enforcement the scheduler cannot provide, so its devices are withheld rather than handed out unenforced.
Risks and Mitigations
Fragmentation (Poisoning)
Risk: A claim with a rare or unique affinity value can lock a high-capacity device to that value, stranding the device’s remaining capacity against peer claims (any priority) that don’t share the value. Pure capacity-stranding, not a priority problem — even claims of equal or lower priority cannot use the locked device. This section covers the peer-claim case only; the priority-aware case is Preemption Cannot Break Affinity Locks .
Mitigation (Alpha): None beyond Filter correctness. The DRA allocator
currently uses a first-fit algorithm with no affinity-aware preference, so
the scheduler does not actively pack compatible claims onto already-locked
devices. Where domain-specific validation is feasible, cluster
administrators can use DeviceClass CEL selectors to restrict which
affinity values are accepted (e.g., constraining subnet IDs to a known
set) — this is an admin-side guardrail against rare or arbitrary values
poisoning devices. Drivers that ship a typed opaque-config CRD can
additionally enforce the same constraints upstream via a validating
admission webhook on the config object, rejecting claims carrying
out-of-policy values before they ever reach the scheduler. Affinity-aware preference (within-node) and a Score
contribution (cross-node) are planned as a beta-scope addition to
DynamicResources.computeScore, following the same per-feature additive
pattern that Prioritized List (KEP-4816, shipped in 1.35) and Extended
Resources (KEP-5004) already use; see Affinity-aware scoring (planned for
Beta)
. Until that lands,
fragmentation mitigation is best-effort and depends on the existing
first-fit ordering of devices in ResourceSlices. General-purpose DRA
scoring discussion continues in
kubernetes/enhancements#4970
,
but KEP-5981’s contribution does not block on a unified framework.
Preemption Cannot Break Affinity Locks
Risk: Standard Kubernetes preemption triggers on resource shortage, not on affinity-lock mismatch. A higher-priority Pod requiring a different affinity value than the current lock cannot evict the lock-holder under standard preemption — capacity is technically free, so preemption is never triggered. This is the same gap that affects all of DRA today, not specific to this KEP.
Mitigation: This KEP does not deliver preemption. The mitigation within this KEP’s scope is scoring/packing — see Fragmentation (Poisoning) and the Beta scoring contribution in Affinity-aware scoring (planned for Beta) . Lock-breaking is a property that emerges when DRA preemption is present in the cluster; see Composition with DRA Preemption (KEP-5690) .
Design Details
API Enhancement
ResourceSlice Spec
type ResourceSliceSpec struct {
// ... existing fields (Driver, NodeName, Pool, Devices, etc.) ...
// SharingAffinity declares the affinity keys that gate sharing
// for every device in this pool, and the CEL expressions that
// extract them from a claim's opaque config. A ResourceSlice may
// set either Devices or SharingAffinity, not both; a pool carries
// its sharing affinity in dedicated metadata slices alongside its
// device-bearing slices. At most 8 affinity keys may be declared
// per slice. See SharingAffinityExtractor for semantics.
//
// +optional
// +listType=map
// +listMapKey=name
// +k8s:maxItems=8
SharingAffinity []SharingAffinityExtractor
}
// SharingAffinityExtractor declares one sharing-affinity key and the
// CEL expression that produces its value from a driver's opaque
// device config. Every entry in a pool's `sharingAffinity` list
// applies to every device in the pool, and a request must satisfy
// every entry for a device in that pool to be viable.
type SharingAffinityExtractor struct {
// Name is the sharing-affinity key name (for example "subnet").
// Must be unique within the pool.
//
// +required
Name string
// Expression is the CEL expression evaluated for this key, over
// the claim's opaque-config object bound as `object`. It must
// return a string. An empty-string return means "no contribution
// for this key from this object" — the idiom an extractor uses to
// ignore opaque-config objects whose apiVersion / kind it does
// not handle.
//
// The length of the expression must be smaller or equal to 10 Ki.
// The cost of evaluating it is also limited based on the
// estimated number of logical steps; the combined cost of all
// extractors in a pool is capped by a shared CEL cost budget.
//
// +required
Expression string
}
const SharingAffinityMaxEntries = 8
// Expression length reuses the existing DRA limit,
// CELSelectorExpressionMaxLength (10 Ki).
// SharingAffinityCELMaxCost is the maximum combined execution cost
// allowed for all sharing-affinity expressions in a single pool,
// mirroring DeviceClaimDerivedAttributeCELMaxCost. Tying the
// collective cost of a pool's extractors to the cost allowed for one
// CEL selector keeps worst-case Filter latency comparable to existing
// device selection.
const SharingAffinityCELMaxCost = 1000000
Scheduler Enhancement
Source of Truth for Affinity Locks
The scheduler derives affinity locks solely from CEL extraction over active claims’ opaque configs — not from device attributes on the ResourceSlice. The driver is NOT required to write locked affinity values back to the ResourceSlice.
- The ResourceSlice declares how to extract sharing-affinity
key/value pairs (
sharingAffinity). - The claims declare what values they need (via their normal opaque
OpaqueDeviceConfiguration.parametersblob — driver’s existing schema). - The scheduler combines these by running CEL at Filter time and
maintains the lock in
AllocatedState.
This avoids two sources of truth that could diverge, eliminates
ResourceSlice churn (no update every time a lock is set/cleared), and
keeps driver implementation simple. Drivers MAY optionally publish
current locked values as regular device attributes for observability
(e.g., visible via kubectl), but the scheduler does not depend on
them.
When the last claim on a device is released, the scheduler clears the
lock. The scheduler’s notion of “clean” is allocation-clean, not
hardware-ready — a device is considered clean once no allocated
claim references it, regardless of whether the driver has finished
in-flight hardware reconfiguration on the node. The driver remains
responsible for device lifecycle: tearing down the old configuration
(via NodeUnprepareResources) and reconfiguring for new claims (via
NodePrepareResources). Driver-level prepare/unprepare sequencing is
the authoritative guard against reuse before reconfiguration
completes; drivers that need stronger guarantees should hold their
own per-device readiness state and reject prepare calls until
reconfiguration is complete.
Safety Model and Responsibility Split
This feature intentionally keeps placement knowledge and hardware enforcement separate:
- Scheduler guarantee: when it has successfully extracted affinity
keys (via the slice’s
sharingAffinityCEL) for all active allocations on a device, it will not intentionally co-place claims with incompatible affinity values on that device. - Conservative fallback: if the scheduler cannot reconstruct the
effective affinity state of a device (for example, due to legacy
claims, opaque configs that no longer produce any non-empty
extraction under the current
sharingAffinityentries, or CEL evaluation failures), it treats that device as unknown and filters it out for new sharing-affinity placements until the device becomes clean. - Driver guarantee: the driver remains the final authority for
programming and validating the actual hardware mode during
NodePrepareResources. - Failure handling: stale scheduler state or races may still cause prepare-time rejection, and that rejection remains the final safety backstop.
Cache Extension: Effective Device State
To prevent race conditions during high-volume scheduling, the scheduler
maintains affinity locks in its internal cache rather than relying on API
server round-trips. This is consistent with how DRA already handles
capacity tracking via inFlightAllocations.
The scheduler’s AllocatedState is extended to track affinity values
alongside consumed capacity:
// AffinityStatus is scheduler-internal cache state, never serialized.
// The zero value is AffinityStatusClean, so a device absent from
// AffinityStates is correctly Clean.
type AffinityStatus int
const (
// AffinityStatusClean: no active claims on the device.
// LockedAffinity is nil.
AffinityStatusClean AffinityStatus = iota
// AffinityStatusLocked: at least one claim is active and the
// scheduler has reconstructed the device's affinity values.
// LockedAffinity holds the current lock.
AffinityStatusLocked
// AffinityStatusUnreconstructable: at least one active claim's
// affinity values cannot be reconstructed (CEL extraction
// failed, did not produce a non-empty value for some declared
// key, or claim predates the feature). The device is filtered
// for new sharing-affinity placements until it becomes Clean.
// LockedAffinity is nil.
AffinityStatusUnreconstructable
)
type AffinityState struct {
// Status is the device's current affinity state. See the
// AffinityStatus constants for the meaning of each value and
// when LockedAffinity is populated.
Status AffinityStatus
// LockedAffinity holds the device's affinity lock when
// Status == AffinityStatusLocked. Nil for Clean and
// Unreconstructable.
LockedAffinity map[string]string
}
type AllocatedState struct {
AllocatedDevices sets.Set[DeviceID]
AllocatedSharedDeviceIDs sets.Set[SharedDeviceID]
AggregatedCapacity ConsumedCapacityCollection
AffinityStates map[DeviceID]AffinityState
}
Filter Phase and Device Selection
Filter phase: For a given node, the scheduler evaluates each device.
A device whose pool declares sharingAffinity (looked up via the
pool’s metadata slice) is a candidate ONLY if:
- It has sufficient consumable capacity (KEP-5075).
- The device’s
AffinityStates[deviceID].Statusis notAffinityStatusUnreconstructable. - For each extractor in the pool’s
sharingAffinitylist, the scheduler evaluates the extractor’s CEL expression against every opaque-config object in scope for this request (per-request scoping viaDeviceClaimConfiguration.Requests). Every CEL expression must evaluate successfully and return a string value. CEL evaluation errors, non-string returns, missing fields, or cost-budget exhaustion are ResourceSlice authoring errors: the scheduler aborts allocation and fails scheduling for the Pod so the problem surfaces immediately, rather than skipping the device (see Extraction failure modes ). - Every extractor is satisfied — each produces a
non-empty return across the request’s opaque configs (Strict
Gating). The contributions from all extractors are merged
into the request’s effective affinity map. Two contributions
producing the same key with different non-empty values are a
ResourceClaim authoring error (self-inconsistent claim); the pod
is marked
UnschedulableAndUnresolvable. Two extractors in the same pool declaring the same key name are rejected before this point — at admission within a slice, and by pool validation across slices (see Key name uniqueness ). - The device’s
AffinityStates[deviceID].LockedAffinityis either empty (unlocked) OR matches the request’s effective affinity map exactly for ALL extracted keys.
If a device has AffinityStates[deviceID].Status == AffinityStatusUnreconstructable, or if not
every extractor is satisfied by the request, or
extraction fails (any reason in #3), the device is not a viable
option for the request. This is the safe default: the driver
declared that sharing requires scheduler-readable extraction, and
a scheduler that cannot reconstruct the current or requested
affinity state cannot evaluate placement safely. Requests that do
not need sharing-constrained devices should target devices in
slices without sharingAffinity.
Device selection within a node: Among feasible devices on a chosen node, alpha does not introduce affinity-aware preference. Device selection continues to use the existing structured-parameters allocator (first-fit). This means that on a node with both a compatibly-locked device with capacity and a clean device, the allocator may pick whichever appears first in the ResourceSlice rather than preferring the locked one.
Affinity-aware preference (within a node) and node-level scoring (across nodes) are planned for beta — see Affinity-aware scoring (planned for Beta) .
Key name uniqueness
Affinity key names must be unique across all extractors in a pool — otherwise the lock state on a device would be ambiguous about which extractor’s value to store. Because every extractor applies to every device in the pool, uniqueness is a property of the key names alone: a set comparison, with no need to consider which devices an extractor applies to. Two enforcement layers cover it:
- Admission-time (exact, within a slice): because
sharingAffinityis a flat list keyed byname(+listType=map +listMapKey=name), duplicate key names within a slice are rejected by generic list validation — no sharing-affinity-specific validation code is required. - Pool validation (across slices): a pool may publish several metadata slices, and the API server validates each slice in isolation, so a collision between extractors in different slices of the same pool cannot be caught at admission. This is the same limitation KEP-4815 has for counter-set references, and it is handled the same way: the allocator validates complete pools and treats a duplicate key name as an invalid pool, failing closed for every device in it rather than silently picking one extractor’s value.
Uniqueness is therefore a static property of the pool, checked before any device is evaluated — not a per-device runtime condition.
Expression limits
Each extractor’s CEL expression is bounded the same way DRA already bounds its other driver-published CEL:
- Length: at most
CELSelectorExpressionMaxLength(10 Ki) per expression — the limit already applied toCELDeviceSelectorandDeviceDerivedAttributeexpressions. - Cost: the estimated evaluation cost of a pool’s extractors is
capped by a shared budget,
SharingAffinityCELMaxCost(1,000,000, roughly 0.1 s), mirroringDeviceClaimDerivedAttributeCELMaxCost. A shared budget rather than a per-expression one keeps the cost of a pool’s whole extractor set comparable to a single CEL selector, which matters because Filter runs every extractor against every in-scope opaque config for every candidate device.
Both limits are validated at admission per slice, and the shared budget is re-checked across a pool’s metadata slices by the same pool validation that enforces key-name uniqueness. Following the existing DRA convention, validation happens only when an expression is set or changed, so changing the limits in a future release leaves already published expressions valid; the scheduler additionally enforces the cost limit at runtime.
Reserve Phase: Tentative Locking
Once a node/device is selected, the Reserve plugin establishes a “tentative lock” in the scheduler cache before the Binding phase. Reserve reuses the key map Filter already extracted — no second CEL pass — and atomically:
- If the device is still unlocked: record the claim’s keys as the device’s
LockedAffinityand proceed. - If the device became locked since Filter (another pod won the race) or
was already locked: re-check the cached keys against the current
LockedAffinity. On match, co-allocate. On mismatch, fail Reserve and let the scheduler retry with the next candidate.
The tentative lock is immediately visible to subsequent scheduling cycles: if Pod-B is evaluated milliseconds after Pod-A’s Reserve (before Pod-A’s bind reaches the API server), Pod-B’s Filter phase sees Pod-A’s tentative lock and either joins it or skips the device.
If scheduling later fails for the pod (Unreserve), the tentative lock is removed unless another already-bound claim is still co-located on the device.
Scheduler Restart: State Reconstruction
On scheduler restart, the in-memory AffinityStates map is empty and must
be rebuilt from already-cached state (bound ResourceClaims and their
ResourceSlices) before the first scheduling cycle. The behavior
contract is:
- Recovery: for each bound claim on a device whose pool declares
sharingAffinity(looked up via the pool’s metadata slice), the scheduler re-derives the claim’s key map and records it as the device’sLockedAffinity. No new API calls are required. - Conservative fallback on ambiguity: if extraction fails, yields no
keys, or yields inconsistent values across claims sharing the device
(which by construction should not happen, but may arise from historical
bugs, manual etcd edits, or version skew), the device is marked
AffinityStatusUnreconstructableand excluded from new sharing-affinity placements until all claims on it drain. The scheduler never infers a lock from ambiguous data. - No new persistence: reconstruction uses the same informer-cached data the scheduler already consumes; no migration, no new API surface.
Feature Gates
This KEP introduces one feature gate:
DRASharingAffinity(alpha): adds thesharingAffinityfield onResourceSlice.spec, theAllocatedState.AffinityStatescache dimension, and the scheduler Filter / Reserve logic that runs CEL extraction over claim opaque configs and matches the result against the lock. The gate must be enabled on bothkube-apiserverandkube-schedulerfor the feature to function. Asymmetric enablement fails closed rather than silently disabling enforcement:apiserver on, scheduler off: drivers can write
sharingAffinityand the apiserver persists it, but the scheduler cannot evaluate it. Rather than allocating those devices with affinity silently unenforced, the scheduler treats every device in a pool that declaressharingAffinityas unallocatable. This mirrors howDRAPartitionableDevicesskips slices usingsharedCountersand devices usingconsumesCounterswhen that gate is off.apiserver off, scheduler on: per standard alpha-field handling, the apiserver strips
sharingAffinityfrom new writes. Slices persisted during a prior enabled period are still served on read, so the scheduler continues to honor existing locks. New attempts to declare affinity do not take effect, but they are not silent: the ResourceSlice controller helper reports aDroppedFieldsErrornamingDRASharingAffinity, the same mechanism that already surfacesDRAPartitionableDevicesandDRADeviceTaints.
Failing closed costs capacity — a scheduler rolled back while pools still declare affinity cannot place new requests on those devices — but ignoring the field is worse: it lets the scheduler co-locate claims the driver declared incompatible, which the driver backstop then rejects at
NodePrepareResources, producing a schedule/reject loop. Operators should have drivers stop publishingsharingAffinitybefore disabling the gate in the scheduler.The expected rollout order mirrors other DRA gates: enable on the apiserver first, then the scheduler; disable in reverse.
Because the claim side carries no new API surface (workloads keep their existing opaque configs), no separate API-only gate is needed.
Examples
ResourceSlice with Sharing Affinity
apiVersion: resource.k8s.io/v1
kind: ResourceSlice
metadata:
name: node1-nics
spec:
driver: networking.example.com
nodeName: node1
sharingAffinity:
- name: subnet
expression: |
object.apiVersion == "networking.example.com/v1"
&& object.kind == "NICConfig"
? object.subnetID : ""
devices:
- name: eth1
allowMultipleAllocations: true
attributes:
networking.example.com/type:
string: "sriov-vf"
capacity:
networking.example.com/slots:
value: "16"
- name: eth2
allowMultipleAllocations: true
attributes:
networking.example.com/type:
string: "sriov-vf"
capacity:
networking.example.com/slots:
value: "16"
The sharingAffinity block declares a single key (subnet) to be
pulled from any opaque-config object that identifies itself as
networking.example.com/v1 / NICConfig. The CEL guard on
object.apiVersion and object.kind ensures the expression returns
"" (no contribution) when applied to an opaque-config object from
a different schema, rather than failing with a missing-field runtime
error. See the Proposal section for the canonical guard idiom and
the multi-version variant.
ResourceClaim (status quo opaque config)
apiVersion: resource.k8s.io/v1
kind: ResourceClaim
metadata:
name: pod-a-nic
spec:
devices:
requests:
- name: nic
exactly:
deviceClassName: shared-nic
config:
- requests: ["nic"]
opaque:
driver: networking.example.com
parameters:
apiVersion: networking.example.com/v1
kind: NICConfig
subnetID: subnet-X # extracted as `subnet` by the slice's CEL
vlanId: 100 # driver-private; not extracted
mtu: 9000 # driver-private; not extracted
Note: The claim shape is unchanged from pre-KEP-5981 DRA — the workload author writes the driver’s existing opaque config. The scheduler runs the slice-declared CEL expression against the parsed
parameters— theapiVersion/kindguard matches, so the expression returnssubnet-Xand the scheduler recordssubnet: subnet-Xfor affinity lock evaluation. The driver-private fields (vlanId,mtu) flow through toNodePrepareResourcesuntouched.
Multi-key Sharing Affinity Example
This example illustrates the alpha semantics when a slice extracts multiple keys.
A driver advertises shared RDMA-capable NICs where both subnet and PKey must match for pods to share the same device:
apiVersion: resource.k8s.io/v1
kind: ResourceSlice
spec:
driver: networking.example.com
nodeName: node1
sharingAffinity:
- name: subnet
expression: |
object.apiVersion == "networking.example.com/v1"
&& object.kind == "NICConfig"
? object.subnetID : ""
- name: pkey
expression: |
object.apiVersion == "networking.example.com/v1"
&& object.kind == "NICConfig"
? object.ibPKey : ""
devices:
- name: mlx5_0
allowMultipleAllocations: true
capacity:
networking.example.com/slots:
value: "16"
A matching claim provides both values inside its driver-defined opaque config:
config:
- requests: ["rdma-nic"]
opaque:
driver: networking.example.com
parameters:
apiVersion: networking.example.com/v1
kind: NICConfig
subnetID: subnet-a
ibPKey: "0x8001"
vlan: "100" # not extracted (no CEL declared for it)
Alpha matching behavior:
- If the device is clean, the first compatible claim sets the lock to:
subnet = subnet-apkey = 0x8001
- A later claim whose extraction yields the same
subnetandpkeymay share the device. - A request whose extraction yields
subnet = subnet-abutpkey = 0x8002makes the device not a viable option because all declared keys must match. - A claim whose opaque config has no
ibPKeyfield will causeobject.ibPKeyto error during CEL evaluation; this is a ResourceSlice authoring error and aborts scheduling for the Pod. - The
vlanfield is ignored because the driver did not declare a CEL expression for it. It flows through opaquely to the driver.
Test Plan
[x] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.
Prerequisite testing updates
Existing DRA scheduling tests should pass before adding sharing affinity tests.
Unit tests
pkg/scheduler/framework/plugins/dynamicresources: Coverage for affinity matching logic, including:- Filter: device with matching lock passes
- Filter: device with conflicting lock is excluded
- Filter: unlocked device with sufficient capacity passes
- Filter: extraction yields no non-empty keys for the claim → device filtered out
- Filter: claim with extra fields in opaque config beyond what the slice’s CEL extracts → extra fields ignored, device passes if extracted keys match
- Filter: claim has multiple opaque configs producing conflicting same-key values across one or more extractors → device filtered out
- Filter: CEL evaluation error (missing field, non-string return, cost budget exhausted) → scheduling aborts with a fatal error, no silent empty lock, no per-device Event spam
- Filter: device with
AffinityStates[deviceID].Status == AffinityStatusUnreconstructableis excluded for new sharing-affinity scheduling - First-fit verification: with two feasible devices on a node (one locked-compatible, one clean), allocator picks the first in ResourceSlice order (no affinity-aware preference; preference lands in beta)
- Reserve: first claim sets lock; second claim with same extracted values succeeds
- Reserve: second claim with conflicting extracted values fails
- Unreserve: tentative lock is rolled back
- Legacy claims with non-reconstructable affinity (no keys produced or CEL fails) cause the device to be marked unknown rather than establishing or joining a lock
- Legacy-claim handling: all scenarios from the
Handling Legacy Claims with Unreconstructable Affinitytable - Compatibility matrix: device in a slice without
sharingAffinityis unaffected — claims with or without matching opaque configs both pass (Legacy/Basic and Permissive Sharing rows) - Strict Gating: device’s pool has
sharingAffinitybut at least one extractor is not satisfied by the claim’s opaque configs (at least one extractor returns empty) → device filtered out - Multi-request scoping: claim with two requests (
mgmt-nicanddata-nic) each with distinct opaque configs → each request resolves independently; one request’s extracted values do not influence the other’s filter decision or lock state on a different device
staging/src/k8s.io/api/resource/v1: Coverage for the newResourceSliceSpec.SharingAffinityfield, including:- Validation:
sharingAffinityexceeding max 8 entries is rejected - Validation: duplicate
namevalues within a slice’ssharingAffinitylist are rejected (declarative+listMapKey=nameuniqueness) - Validation: CEL expressions are syntactically valid at admission (parse-check only; runtime cost is bounded per-eval at the scheduler)
- Validation: a single
ResourceSlicewith bothdevicesandsharingAffinityset is rejected (mutual exclusion) - Round-trip serialization of
SharingAffinityExtractor
- Validation:
Integration tests
- Affinity matching with multiple claims to same device: in a single-device topology, verify that a second compatible claim shares the locked device and extends the affinity lock. Tests Filter correctness and Reserve-phase lock extension; preference for the locked device when alternatives exist is a beta concern (see Affinity-aware scoring (planned for Beta) ).
- Affinity mismatch causing allocation to different device
- Affinity lock clearing when all claims release a device
- Interaction with consumable capacity constraints (KEP-5075)
- Scheduler restart:
AffinityStatescorrectly reconstructed by re-running CEL extraction against the opaque configs of existing bound ResourceClaims; devices with non-reconstructable active claims (no keys produced, CEL eval failure) haveAffinityStates[deviceID].Status == AffinityStatusUnreconstructable - Parallel scheduling: two Pods whose extracted affinity values conflict target the same device — one wins Reserve, the other is requeued
DRASharingAffinitydisabled in the scheduler: devices in a pool that declaressharingAffinityare skipped, not treated as unconditionally shareable; pools without the field are unaffectedDRASharingAffinitytoggled: enabling after claims exist does not disrupt already-bound workloads, and legacy in-use devices are conservatively filtered until clean- Invalid opaque config at scheduling time: regression test that malformed configs (no apiVersion/kind, unparseable parameters) cause CEL extraction to fail deterministically and the device is filtered out rather than crashing the scheduler
- Permissive Sharing (slice has no AE): Slice without
sharingAffinity, claim with arbitrary opaque config — verify scheduler allows the allocation and opaque config is not evaluated for affinity - Ghost Lock: Pod is Assumed (tentative lock set) but Bind fails — verify the lock is cleared immediately and the next Pod in the queue can claim the device with a different affinity value
- Legacy Device Migration: 5 Pods are already running on NICs in a slice
without
sharingAffinity; the driver republishes the slice withsharingAffinitydeclared; a 6th Pod arrives with a matching opaque config — verify each affected device hasAffinityStates[deviceID].Status == AffinityStatusUnreconstructableand the 6th Pod is filtered from those devices until all legacy claims drain - Partial Key: Pool declares two extractors,
subnetandpkey. Claim’s opaque config hassubnetID(sosubnetresolves) but noibPKey(so thepkeyextractor returns empty and is not satisfied) — verify the device is filtered out - Multiple Extractors, All Required: Pool declares two extractors,
vendorandsubnet. A claim providing bothvendorandsubnetID— verify the device is eligible and the merged effective affinity map is{vendor, subnet}. The same claim missingvendorleaves the first extractor unsatisfied — verify the device is filtered out, confirming that satisfaction is required across every extractor in the pool, not just one. - Duplicate Key Across Metadata Slices (invalid pool): A pool
publishes two metadata slices, each declaring an extractor with
key
subnet. Neither slice is rejectable at admission in isolation — verify the allocator’s pool validation marks the pool invalid and fails closed for every device in it, rather than silently picking one extractor’s value. After the driver republishes with the colliding key renamed, the next scheduling cycle succeeds. - Expression limits: an expression one byte over
CELSelectorExpressionMaxLength(10 Ki) is rejected at admission and one at exactly the limit is accepted; a set of extractors whose combined estimated cost exceedsSharingAffinityCELMaxCostis rejected, both within one slice and across a pool’s metadata slices - First-Fit Behavior: Two devices available on a node, one already locked to subnet-X, one clean; new claim whose extraction yields subnet-X — verify the claim is allocated successfully (Filter excludes nothing; either device is feasible) and that the chosen device matches the allocator’s existing first-fit ordering. Alpha does not require the locked device to be preferred; affinity-aware preference is delivered in beta
- Driver Backstop: Slice has no
sharingAffinity, two claims with incompatible opaque config land on the same device — verify scheduler allows both (permissive), andNodePrepareResourcesrejects the incompatible claim - NodePrepareResources failure does not clear lock: Claim is bound and
lock is set in the scheduler cache, but
NodePrepareResourcesfails on the node — verify the affinity lock remains in the scheduler cache - Multi-request, multi-device, no cross-talk: Single claim with two
requests (
mgmt-nic,data-nic), each with a distinct opaque config yieldingsubnet=Aandsubnet=Btargeting two different devices in slices that declare extraction — verify both requests succeed in the same scheduling cycle and each device locks to its own request’s extracted values without influence from the sibling request - Restart with inconsistent reconstructable locks: two active
reconstructable claims on the same device produce conflicting extracted
key values during restart reconstruction — verify the device is marked
AffinityStates[deviceID].Status = AffinityStatusUnreconstructable, a warning is logged, and new sharing-affinity scheduling is blocked on that device until all claims drain sharingAffinitymutation with active claims: slice republishes with renamed CEL keys or a new extractor; pre-existing claims now extract a different key set than the currentLockedAffinity— verify the device hasAffinityStates[deviceID].Status == AffinityStatusUnreconstructableafter the next reconciliation/restart and is filtered out for new sharing-affinity scheduling until existing claims drain- CEL cost budget: pathological CEL expression intentionally exhausts the per-eval cost limit — verify the scheduler treats this as an extraction failure (device filtered out + Event), does not stall, and remains responsive to other scheduling work
e2e tests
- End-to-end test with mock DRA driver publishing
sharingAffinity - Multi-pod scheduling: Pods with matching extracted affinity values share the same device
- Multi-pod scheduling: Pods with conflicting extracted affinity values are placed on different devices
- Lock lifecycle: last Pod deleted → lock cleared → new Pod with different affinity value can claim the device
- Rollout scenario: existing Pods running on devices in a slice with no
sharingAffinity; driver republishes the slice withsharingAffinity; verify existing Pods continue running and new Pods respect the new constraint after legacy claims drain
Graduation Criteria
Alpha
- Feature implemented behind a single feature gate
DRASharingAffinitycovering both theResourceSlice.spec.sharingAffinityAPI field and the scheduler logic that runs CEL extraction and matches results against the lock - API field added to ResourceSlice (
SharingAffinityExtractoronResourceSliceSpec) - Scheduler runs CEL extraction over claim opaque configs to derive affinity keys; CEL evaluates in the standard ValidatingAdmissionPolicy environment with a per-evaluation cost budget
- Scheduler Filter plugin enforces affinity matching
- Scheduler tracks affinity in
AllocatedState.AffinityStates - Unit and integration tests
- Documentation for driver authors covering CEL expression authoring,
multi-version guard idioms (e.g.,
object.apiVersion == "x/v1" && object.kind == "Foo" ? ... : ""), and rollout guidance - Alpha documentation explicitly calls out the lack of lock-breaking preemption semantics for incompatible locks
- Alpha documentation explicitly calls out string-only affinity matching
(alpha CEL expressions must return
string) - Distinct alpha diagnostics emitted via scheduler logs and (best-effort) events for: (a) compatibility mismatch with the current lock, (b) no keys produced by any extractor for the claim, (c) CEL evaluation failure (including missing fields, non-string return, or cost-budget exhaustion), and (d) unknown lock state due to legacy / non-reconstructable / inconsistent active claims — sufficient to attribute a filtered scheduling decision without relying on metric labels alone
Beta
- Gather feedback from DRA driver developers
- Address any issues found in alpha
- Affinity-aware scoring contribution: Extend
DynamicResources.computeScorewith a sharing-affinity term — prefer nodes where the target device is already locked to a compatible value (consolidation), then clean devices, with unknown-affinity devices deprioritized. Extend the structured-parameters allocator’s per-node device selection with the same preference. Follows the per-feature additive scoring pattern that Prioritized List (KEP-4816, shipped in 1.35) and Extended Resources (KEP-5004) already use; not blocked on the general-purpose scoring discussion in #4970 . See Affinity-aware scoring (planned for Beta) . - Observability of lock state: Surface effective per-device lock state
for operators — exact mechanism TBD in beta. Candidates include
scheduler-side metrics/gauges keyed by device and parameter hash,
scheduler Events on Pending pods naming the locking claim, aggregation
tooling over existing
ResourceClaim.status.allocation, or a dedicated scheduler-owned API resource. - E2e tests stable
- Performance validation with high pod churn
GA
- At least 2 production drivers using sharing affinity
- No significant issues reported
- Conformance tests if applicable
Upgrade / Downgrade Strategy
Upgrade: Existing ResourceSlices without sharingAffinity
continue to work. New field is additive. See the
Compatibility Matrix
for how the scheduler and
driver behave across all combinations of slice sharingAffinity
and claim opaque config presence.
Recommended Rollout
This is a single-stage rollout: no claim-side API change is required, and workloads keep their existing opaque configs throughout.
The recommended sequence:
Enable the gate (
DRASharingAffinity) on apiserver and scheduler. Until any slice declaressharingAffinity, this is a no-op for scheduling.Publish
sharingAffinityon ResourceSlices: add the CEL extractor block to the relevant slices. Strongly recommended to do this on idle slices first (see “AddingsharingAffinityto an in-use slice” below). From this point on, the scheduler locks a device when a claim targeting it produces a non-empty CEL extraction; for requests whose opaque configs produce no non-empty extractions, the device is not a viable option.
Why this order: when the DRASharingAffinity gate is OFF on the
apiserver, sharingAffinity is stripped from incoming ResourceSlice
writes (standard alpha-field handling). A driver that publishes the
field before the gate is enabled will see it dropped at write time —
reported by the ResourceSlice controller helper as a
DroppedFieldsError, not silently. Enabling the gate later does not
retroactively restore it, and the driver must republish.
Additional rollout consideration — workload-schema readiness:
independent of step ordering, enabling extraction on slices before
workloads have rolled to a compatible opaque-config schema may make
the affected devices no longer viable options for those workloads on
their next scheduling attempt, either landing them on non-extraction
slices (capacity strand) or driving them to Pending if no
non-extraction slices exist. The slice update should be coordinated
with any required workload-side opaque-config evolution.
For the common case where the driver’s opaque config schema is unchanged and the slice’s CEL expressions simply pull existing fields out, no workload-side change is needed at all.
Minimizing capacity stranding: when adding sharingAffinity to
slices for the first time, the ideal sequence is:
- Wait for affected devices to be idle (clean).
- Update the ResourceSlice to include the
sharingAffinityblock. - Allow the scheduler to establish the first known lock with a new claim.
During mixed rollouts (some slices with sharingAffinity, some
without), Strict Gating automatically routes claims whose opaque
configs lack the declared keys onto non-extraction slices —
extraction-slice devices are not viable options for those requests
by design. The remaining gap is for
claims whose opaque configs do produce the keys: in alpha they pass
Filter on both extraction and non-extraction slices and the allocator
has no preference between them, so capacity can strand on either side
of the rollout. Affinity-aware preference is delivered in beta. For
predictable rollout behavior, drain devices before adding
sharingAffinity, as recommended above.
Adding sharingAffinity to an in-use slice: A driver may add or
update sharingAffinity on a slice whose devices already have bound
ResourceClaims. The scheduler handles this conservatively:
- Pre-existing claims continue to run and are not evicted.
- For each in-use device on the updated slice, the scheduler re-runs
CEL extraction against every bound claim’s opaque config and
populates
LockedAffinityif all claims yield a consistent key map. If extraction fails or claims yield conflicting maps, the device is markedAffinityStatusUnreconstructableand excluded from new sharing-affinity placements until all its claims drain. - Devices on the slice with no bound claims are unaffected: they behave normally on next allocation.
Driver Upgrades and Schema Evolution: The opaque-config schema and the slice’s CEL extractors both evolve with driver releases. The recommended playbook depends on the kind of change:
| Upgrade class | Recommended action |
|---|---|
| Driver patch with no schema change, same CEL | None. Driver republishes an identical slice; transparent to the scheduler. |
| Driver release with additive schema fields, same extracted keys | None. Existing CEL guards on apiVersion/kind continue to match; new fields are unread by the extractor and flow through to the driver. |
| Driver release adds v2 schema alongside v1 | Extend each extractor’s CEL with a v1 OR v2 apiVersion guard reading the same field path (the canonical multi-version idiom). Both old and new workloads extract identically. After workloads have migrated off v1, a later driver release drops the v1 guard. |
Driver release with breaking schema rename (e.g., subnetID → subnet) — affects opaque-config field path only; declared key set unchanged | Republish the slice with an extractor that reads both paths into the same key, e.g. has(object.subnet) ? object.subnet : object.subnetID. Both old and new workloads extract identically; the legacy branch can be dropped in a later driver release once workloads have migrated, or kept indefinitely (the cost is one ternary in the CEL string). Existing locks remain valid — no Unreconstructable transitions. |
Driver release with renamed extracted key (e.g., subnet → vpc-subnet) — affects the declared key set itself | This is a sharingAffinity mutation: re-extraction against bound claims yields a different key map than the stored LockedAffinity. Devices with active legacy claims transition to AffinityStates[deviceID].Status == AffinityStatusUnreconstructable until they drain. Drivers should prefer to drain affected slices first (per the rollout sequence above) and stage the change to off-hours. |
| DaemonSet rolling upgrade across nodes | Each node’s slice flips when its driver pod restarts. Transient heterogeneity across slices is bounded by the rollout window. Strict Gating plus Unreconstructable quarantine make the window safe but capacity-stranding; rolling-update maxSurge / maxUnavailable should be tuned with this in mind. |
| Workload-side schema lag (workloads behind driver) | Drivers should keep multi-version-guard CEL through at least one workload rollout cycle. Otherwise workloads hit “no keys produced → Strict Gating → Pending” until they upgrade. |
| Driver pod removed or unhealthy | Existing DRA behavior applies — the slice eventually becomes stale and is garbage-collected by the kubelet plugin manager. Devices in a stale slice are unschedulable regardless of sharingAffinity. No additional handling is introduced by this KEP. |
Handling requests that produce no keys: The scheduler treats a slice with
sharingAffinity as a protected resource. If a request’s opaque
configs yield no non-empty value from any declared extractor’s CEL,
no device in that slice is a viable option for the request. CEL evaluation
errors (missing field, non-string return, cost-budget exhaustion) are
treated the same way. API validation prevents most malformed CEL
expressions from reaching the scheduler in the first place (parse-check
at admission).
Downgrade: If the feature gate is disabled:
- The apiserver strips the
sharingAffinityfield from writes that CREATE new ResourceSlices, so a slice freshly created with the gate off has nosharingAffinity. - The apiserver PRESERVES the
sharingAffinityfield on writes that UPDATE existing ResourceSlices where the field was already set on the prior object (ratcheting drop-on-disable). This prevents silent data loss across gate flips. Users may explicitly clear the field by setting it to null. - The scheduler skips every device in a pool that still carries
sharingAffinity, rather than returning those pools to unconditional sharing. Drivers clear the field to make them allocatable again. - The DRA driver becomes the sole authority for enforcing hardware
compatibility at
NodePrepareResourcesonce the field is cleared.
Version Skew Strategy
- kube-apiserver: Must be upgraded first to accept the new
sharingAffinityfield onResourceSlice. - kube-scheduler:
- A scheduler that understands this feature enforces extraction,
tracks
AffinityStates, and may conservatively setAffinityStates[deviceID].Status = AffinityStatusUnreconstructablewhen effective affinity cannot be reconstructed. - A scheduler that has the code but runs with the gate off skips
every device in a pool that declares
sharingAffinityrather than allocating it unenforced. - A scheduler predating the feature does not recognize the field at
all and cannot fail closed, so placement may be overly permissive
and the DRA driver remains the final safety backstop during
NodePrepareResources. This is the case the apiserver-first rollout order exists to keep short.
- A scheduler that understands this feature enforces extraction,
tracks
- kubelet: No changes required; kubelet does not interpret
sharingAffinity. - DRA driver:
- Drivers publish ResourceSlices with the
sharingAffinityfield. - Drivers must continue validating actual hardware compatibility at prepare time, especially during skew where an older scheduler may not enforce affinity constraints.
- Drivers publish ResourceSlices with the
During version skew, the main outcomes are permissive scheduling by an older scheduler or conservative filtering by a newer scheduler when affinity state cannot be reconstructed. Both are operationally safe as long as the driver continues rejecting incompatible prepare-time configurations.
Production Readiness Review Questionnaire
Feature Enablement and Rollback
How can this feature be enabled / disabled in a live cluster?
- Feature gate
- Feature gate name:
DRASharingAffinity— gatesResourceSlice.spec.sharingAffinityand the scheduler logic that runs CEL extraction and enforces matching - Components depending on the feature gate: kube-apiserver, kube-scheduler
- Feature gate name:
Does enabling the feature change any default behavior?
Not on its own. The feature gate is a precondition; behavior only
changes for slices that explicitly opt in via the new
sharingAffinity field, and slices without it behave exactly as
before. Once a driver publishes slices that opt in, however,
DeviceRequests evaluated against those slices will be subject to
CEL extraction and affinity-lock matching at scheduling time without
any change to the claim itself — that is the intended effect of the
driver’s opt-in, but it means a cluster can see scheduling outcomes
shift purely from a slice-side change. With the gate enabled, the timing of that
shift is controlled by when drivers begin advertising
sharingAffinity on their slices; with the gate disabled, the
apiserver strips the field on writes that create new slices. Slices
that already have sharingAffinity set when the gate flips off keep
their data (ratcheting drop-on-disable — see the rollback question below).
Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?
Yes. Disabling DRASharingAffinity causes:
- API server to strip the
sharingAffinityfield on writes that CREATE new ResourceSlices (writes succeed, field is not stored). - API server to PRESERVE the
sharingAffinityfield on writes that UPDATE slices where the field was already set on the prior object (ratcheting drop-on-disable; prevents silent data loss across gate flips). Users may explicitly clear the field by setting it to null. - Scheduler to skip every device in a pool that still declares
sharingAffinity, rather than allocating it with enforcement off (see Feature Gates ).
Existing allocations continue to work. New allocations onto
affinity-declaring pools stop until drivers republish those pools
without sharingAffinity, so a complete rollback is: disable the gate
in the scheduler, then have drivers clear the field.
What happens if we reenable the feature if it was previously rolled back?
The scheduler resumes enforcing sharingAffinity for future placement
decisions. Existing allocations are not evicted.
ResourceSlices that ALREADY HAD sharingAffinity set when the gate was
disabled retain the field (ratcheting drop-on-disable preserved the
data across the gate flip). The scheduler picks up enforcement on those
slices immediately. ResourceSlices that were newly CREATED while the
gate was off will not have the field; drivers may republish those
specific slices with sharingAffinity to opt them in.
While the gate was disabled the scheduler skipped affinity-declaring
pools entirely, so it introduced no inconsistent co-location there.
Devices can still need reconstruction: slices CREATED while the gate
was off had the field stripped, so their pools were allocated
permissively, and claims predating the feature carry no extractable
keys. On reenable, the scheduler reconstructs lock
state per device by re-running CEL extraction over each device’s active
claims. A device is marked AffinityStatusUnreconstructable if the
extracted key maps across its claims are inconsistent, or if extraction
produces no non-empty keys at all. Such devices are excluded from new
sharing-affinity placements until all active claims drain, after which
they become clean and can establish a lock normally.
Are there any tests for feature enablement/disablement?
Yes, unit tests will cover the feature gate behavior for API validation and scheduler logic.
Rollout, Upgrade and Rollback Planning
How can a rollout or rollback fail? Can it impact already running workloads?
Rollout failure modes include:
- Older scheduler after API enablement: a scheduler without the
gate skips every device in a pool that declares
sharingAffinity, so those devices are unschedulable until the scheduler is upgraded or drivers stop publishing the field. - Newer scheduler enabling conservative handling on legacy in-use devices: devices with non-reconstructable active claims may be filtered until they are clean, which can temporarily reduce effective schedulable capacity.
Rollback failure mode: if the scheduler is rolled back while the API
server still serves the field, devices in affinity-declaring pools stop
accepting new allocations until drivers clear sharingAffinity.
Running workloads are not evicted by this feature; the impact is on future placement decisions, not on already-running pods.
What specific metrics should inform a rollback?
Each signal below names a metric trajectory and the condition under
which rollback beats a forward fix. Two levers exist and they are not
interchangeable. Having drivers republish the affected pools without
sharingAffinity is the per-pool lever: it restores permissive sharing
immediately and needs no scheduler restart. Disabling the gate in the
scheduler is the cluster-wide lever, but because a scheduler with the
gate off skips affinity-declaring pools entirely, it must be paired
with drivers clearing the field or those devices become unschedulable.
Unless stated otherwise, “rollback” below means the driver-side lever.
sharing_affinity_unreconstructable_devicesdoes not decay over an extended window (remains near its post-enablement peak after the expected workload-churn window). Indicates legacy in-use devices are not draining naturally and conservative handling will continue to suppress schedulable capacity indefinitely. Rollback restores permissive sharing for these devices; forward-fix has no lever.sharing_affinity_lock_conflict_totalrises for claims that operators expect to be compatible (cross-check againstsharing_affinity_compatible_reuse_totalstaying flat). Indicates the driver’s CEL extractor is producing incorrect keys and over-filtering legitimate placements. Rollback re-enables fungible sharing while the extractor is corrected and re-published.sharing_affinity_no_keys_extracted_totalrises for workloads that previously placed successfully. Indicates extraction is silently dropping required parameters — typically claim opaque-config schema drift or extractor mis-targeting. Rollback restores placement while the driver re-publishes correct extractors.- Sustained increase in unschedulable DRA-backed pods beyond pre-enablement baseline (post-enablement P95 stays elevated over the cluster’s typical workload-churn window, with no corresponding workload increase). Indicates the combined effect of conservative handling and extractor-driven filtering is starving production placements faster than the feature’s benefits accrue.
- Scheduler Filter/PreFilter latency regression measurable in
scheduler_framework_extension_point_duration_secondstraceable to the affinity code path on DRA-heavy clusters. Rollback removes the per-claim CEL evaluation cost while the regression is profiled and fixed. - Scheduler panics or crashes in the affinity code path that cannot be patched in a short window. Rollback (disable the feature gate) is the lower-risk remediation than rolling an in-flight fix to production schedulers.
Note: rising rates of driver prepare-time rejections are not a rollback signal for this feature — those are exactly the failure mode KEP-5981 reduces. If they rise after enablement, the cause is upstream of the scheduler (driver bug, extractor bug) and rolling back would make them more frequent, not fewer.
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
Will be tested before beta.
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?
No.
Monitoring Requirements
How can an operator determine if the feature is in use by workloads?
Two signals, in order of cost-to-collect:
Metrics (recommended for fleet-scale visibility).
sharing_affinity_locked_devicesreports how many devices the scheduler currently has under affinity lock. Any non-zero value confirms the feature is in active use in this cluster — the primary fleet-adoption signal.- A rising
sharing_affinity_extractor_evaluations_totalcounter indicates the scheduler is actively running CEL extraction against publishedsharingAffinityslices, even if no locks have stuck yet.
Both metrics are defined in the “Are there any missing metrics…” section below.
Object inspection (single-cluster diagnosis).
ResourceSliceobjects that declaresharingAffinityshow driver-side opt-in;ResourceClaims whose opaque configs yield non-empty keys under one of the declared extractors show workload-side participation. Useful for diagnosing a specific cluster, but not for fleet observability.
How can someone using this feature know that it is working for their instance?
There are no events or logs to show that the scheduler is excluding devices due to affinity locks, and adding them may be noisy. Instead, users can observe that the affinity is respected. A user should be able to observe that:
- compatible claims are eligible to reuse already-locked devices (alpha does not actively prefer them; affinity-aware preference lands in beta),
- when the scheduler has reconstructable affinity state, devices with incompatible affinity locks are not viable options for incompatible requests — gating happens before bind/prepare,
- devices with unknown legacy affinity state are conservatively excluded until they become clean.
In practice, this should be visible through scheduler logs, scheduler events, and (where implemented) scheduler metrics.
What are the reasonable SLOs (Service Level Objectives) for the enhancement?
This enhancement should not materially regress baseline DRA scheduling latency
for clusters that do not use sharingAffinity.
For clusters that do use the feature, the primary objective is correctness of compatibility-aware placement with bounded incremental scheduling overhead.
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
Useful SLIs include:
- rate of scheduling attempts filtered due to sharing-affinity mismatch (no keys produced, CEL eval error, or lock mismatch),
- rate of devices with
AffinityStates[deviceID].Status == AffinityStatusUnreconstructable, - share of successful placements that reuse already-locked compatible devices,
- prepare-time rejections by the DRA driver caused by incompatible hardware configuration.
Are there any missing metrics that would be useful to have to improve observability of this feature?
This feature would benefit from scheduler-observable counters, gauges, and/or events for:
sharing_affinity_locked_devices— number of devices currently holding at least one active affinity lock. The primary “feature in active use” signal: non-zero means the scheduler is actively maintaining affinity constraints in this cluster.sharing_affinity_unreconstructable_devices— number of devices currently in conservative quarantine because their lock state cannot be reconstructed (legacy active claims, or extractors that no longer produce keys for those claims). If paired withsharing_affinity_locked_devices, they describe the cluster’s current device-state distribution under the feature.sharing_affinity_extractor_evaluations_total— total CEL extractor evaluations the scheduler has performed. Useful as a “feature is being exercised” signal independent of placement outcomes. If this is rising butsharing_affinity_locked_devicesstays at zero, the scheduler is evaluating extractors but no new locks are being established. Common causes, distinguishable by inspecting the other metrics in this list:- candidate devices are unreconstructable (
sharing_affinity_unreconstructable_devices > 0) — legacy claims have not drained yet, - extractors do not match the claims’ opaque-config schemas (
sharing_affinity_no_keys_extracted_totalrising) — driver configuration or claim-template issue, - lock conflicts are blocking placements (
sharing_affinity_lock_conflict_totalrising) — competing claims with incompatible affinity values.
- candidate devices are unreconstructable (
sharing_affinity_lock_conflict_total— candidate devices filtered out during Filter because the device’s existing affinity lock is incompatible with the claim’s extracted keys. This is the “feature is gating incompatible placements as designed” signal — non-zero is expected and healthy; the rate scales with claim diversity, not feature health.sharing_affinity_no_keys_extracted_total— candidate devices filtered out during Filter because the slice’s extractors produced an empty effective affinity-key map for the claim. The adoption-gap signal: a high ratio ofno_keys_extracted_total / extractor_evaluations_totalmeans the feature is wired up but workloads aren’t engaging. Settles low in mature clusters; drops over time as workload teams update ClaimTemplates during rollout.sharing_affinity_compatible_reuse_total— successful placements onto an already-locked device whose affinity values matched. It can stay zero when the feature is in use but the scheduler picks a different device for scoring reasons (spread, topology preference). Usesharing_affinity_locked_devicesas the in-use indicator andsharing_affinity_compatible_reuse_totalas the “preference is delivering value” indicator.
Counters above give cluster-wide visibility; individual pods that fail to schedule also need a clear per-pod reason. The scheduler should surface a specific diagnostic on the pod’s PodScheduled=False condition (and corresponding scheduling event) whenever an affinity check rejects a device, so the workload owner can tell why their pod is stuck. For example:
- claim’s opaque configs produced no non-empty key under any
sharingAffinityextractor on slice<slice>(device<id>filtered), - CEL expression
<key>failed to evaluate against claim’s opaque config:<error>(device<id>filtered), - device
<id>is locked to incompatible affinity values, - device
<id>has unknown affinity state due to legacy or invalid active claims.
Dependencies
Does this feature depend on any specific services running in the cluster?
- KEP-5075 (Consumable Capacity) for multi-allocatable devices
Scalability
Will enabling / using this feature result in any new API calls?
No new API calls. Affinity data is extracted from ResourceSlice and ResourceClaim objects already fetched by existing informers.
Will enabling / using this feature result in introducing new API types?
No. Only new fields on existing types.
Will enabling / using this feature result in any new calls to the cloud provider?
No.
Will enabling / using this feature result in increasing size or count of the existing API objects?
- ResourceSlice: One or more additional metadata slices per pool
that uses this feature (each carrying
sharingAffinityand zero devices); the metadata slices in a pool jointly hold up to 8 extractors, each with up to 8 CEL expressions. Device-bearing slices are unchanged. - ResourceClaim: No change. The claim side carries no new fields.
Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?
Negligible. The Filter phase evaluates every extractor’s CEL against
every opaque-config object on the claim. Worst-case CEL evaluation count per candidate device is
O(extractors × opaque-configs), bounded by ≤8 extractors × N
opaque configs (typically ≤2), and gated by
the shared per-pool cost budget (SharingAffinityCELMaxCost, equal to
the cost allowed for a single CEL selector). CEL programs are
compiled at slice-write admission time and cached; per-evaluation cost
is small and bounded. The lock-comparison itself is O(k) where k ≤ 8 — a
map lookup per key.
Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?
No. The per-component impact is bounded:
- Scheduler RAM:
AffinityStatesadds onemap[string]string(up to 8 entries) per device with active affinity locks — proportional to active shared allocations, not total devices. A cluster with 1,000 actively-shared devices carries well under 1 MB of state; unshared devices contribute zero. Compiled CEL programs are cached cluster-wide keyed by expression string; the cache is bounded by ≤8 extractors per slice, with cross-slice deduplication in practice. - Scheduler CPU: Running CEL extraction during Filter is small per-candidate and bounded by the cost budget. Programs are compiled once and reused.
- etcd disk: Slightly larger ResourceSlice objects (bounded by the caps above). ResourceClaim size is unchanged.
Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?
No.
Troubleshooting
How does this feature react if the API server and/or etcd is unavailable?
Like existing scheduler-driven DRA logic, this feature depends on informer state and cached API data. Temporary API server or etcd unavailability does not by itself invalidate already-computed in-memory lock state, but new pods will not be scheduled during unavailability. Sustained control-plane unavailability may delay reconciliation of claim release, slice updates, or restart reconstruction.
The driver remains the final enforcement authority at prepare time.
What are other known failure modes?
Known failure modes include:
- No keys produced: the request’s opaque configs do not yield any
non-empty value from at least one
sharingAffinityextractor on the slice; the device is not a viable option for the request (Strict Gating outcome). - CEL evaluation failure: a CEL expression returns a non-string, references a missing field on the parsed opaque parameters, or exhausts the per-evaluation cost budget. ResourceSlice authoring error: the scheduler aborts allocation and fails scheduling for the Pod. The scheduler never silently establishes an empty lock.
- Unreconstructable affinity state: the device has active allocations whose affinity cannot be reconstructed (legacy claims, extractors changed such that they no longer produce keys, or CEL failure during reconstruction), so it is conservatively filtered until clean.
- Prepare-time driver rejection: despite scheduler filtering, the driver may still reject an incompatible or stale placement and that rejection is the final safety backstop.
- Partial feature gate enablement: if the feature gate is enabled
on the API server but not the scheduler (or vice versa), the
sharingAffinityfield may be persisted but not enforced, or enforced from cached objects but unable to be persisted on new writes. Ensure the gate is enabled on bothkube-apiserverandkube-scheduler.
What steps should be taken if SLOs are not being met to determine the problem?
Recommended debugging flow:
- Inspect the relevant
ResourceSliceand confirm it declares the expectedsharingAffinityextractors (well-formed CEL expressions, expected key names). - Inspect the
ResourceClaimand confirm at least one of its opaque-config objects has the shape the slice’s CEL expressions expect (matchingapiVersion/kindguards, populated fields). - Mentally (or via
kubectl+ a CEL test harness) run each CEL expression against the claim’s parsed opaque parameters and verify it returns a non-empty string value. - Check whether the target device is already locked to incompatible values (lock state is in the scheduler’s in-memory cache — check scheduler logs for filter reasons mentioning affinity mismatch).
- Check whether the device is being treated as having unknown affinity state because of legacy or non-reconstructable active claims.
- Review scheduler logs/events for explicit filter reasons (no keys produced, CEL eval failure, cost-budget exhaustion, lock mismatch, unknown state).
- If the scheduler allowed placement but the driver rejected prepare, inspect driver logs to determine whether the issue was stale scheduler state, unsupported config, or an actual device-level incompatibility.
Implementation History
2026-03-27: Initial KEP issue created
2026-03-30: KEP document drafted
2026-04-28: Pivoted from a well-known JSON schema inside
OpaqueDeviceConfigurationto a typedStructuredsibling field onDeviceConfiguration, based on wg-device-management feedback that the scheduler should not interpret opaque payloads.2026-05-14: Pivoted again from the typed
Structuredsibling onDeviceConfigurationto driver-published CEL extraction onResourceSlice.spec.sharingAffinity, per @johnbelamaric and @pohly review feedback on PR #5987. The claim side reverts to the status-quo opaque config (no new typed claim-side surface), and the scheduler’s only assumption becomes “find some GVK objects to which it can apply CEL.” Removed theDRAStructuredDeviceConfigurationfeature gate (no longer needed). Consolidated into a singleDRASharingAffinitygate.2026-05-16: Per @pohly review feedback (#discussion_r3246775761), dropped the
GVKfield fromSharingAffinityExtractor. GVK matching is now the CEL author’s responsibility: a CEL expression guards onobject.apiVersionandobject.kindand returns the empty string when it does not apply to a given opaque-config object. The scheduler unions non-empty returns across all (extractor, opaque-config) pairs. Simplifies the API and lets one extractor handle multiple compatible config versions (e.g., v1beta1 + v1 with same field paths) without duplication.2026-05-17: Added a Driver Upgrades and Schema Evolution playbook under the Recommended Rollout section, covering schema-additive, multi-version-add, breaking-rename, key-rename, DaemonSet rolling upgrade, workload-side lag, and driver-pod removal cases. Codifies the multi-version-guard CEL idiom as the recommended upgrade primitive for breaking schema changes.
2026-09-28: Addressed @pohly review feedback on PR #5987:
- Replaced each entry’s
cel map[string]stringwith a flatSharingAffinityExtractor{Name, Expression}list keyed byname(#discussion_r4120271856, #discussion_r3199293328). - Dropped the per-extractor
Selector; pools usingsharingAffinityare homogeneous, and heterogeneous pools became a rejected alternative (#discussion_r4120408066). - ResourceSlice authoring errors now abort allocation and fail scheduling for the Pod instead of skipping the device and emitting an Event (#discussion_r4120337098, #discussion_r4120605895).
- A scheduler with the gate off now skips every device in a pool
declaring
sharingAffinityrather than ignoring the field (#discussion_r4120708639, #discussion_r4120870443). - Bounded each CEL expression by length and a shared per-pool cost budget — see Expression limits (#discussion_r4120767010).
AffinityStatusis now an int enum rather than a string; it is scheduler-internal cache state, and the zero value isAffinityStatusClean(#discussion_r4120810687).
- Replaced each entry’s
Drawbacks
- Adds a new cache dimension (
AffinityStates) to the scheduler’s allocation tracking, increasing the surface area for reconstruction bugs on restart - Once a device is locked, its effective affinity cannot change until all claims on that device are released
- Fragmentation risk remains if affinity values are too fine-grained
- Conservative handling of legacy in-use devices can temporarily strand schedulable capacity during rollout or migration
- If a pool declares
sharingAffinitybut a request’s opaque configs do not satisfy every extractor for a given candidate device (every extractor returning non-empty), that device is not a viable option for that request under the “Strict Gating” rule; if no device in the pool has all extractors satisfied by the request, the entire pool is filtered out. Drivers should coordinate with workload teams to ensure claims carry compatible opaque configs matching the published extractors before enablingsharingAffinityon the pool. - Per-evaluation CEL cost is bounded but non-zero; under pathological CEL expressions a slice could elevate per-candidate Filter cost. The per-eval cost budget and admission-time parse check are the primary mitigations.
- Affinity locks are purely in-memory with no API or status field to
inspect which devices are locked to which values. Debugging lock
state in alpha requires scheduler logs; a future enhancement (tracked
under Beta graduation) is to surface effective lock state via
scheduler-side metrics, scheduler Events on Pending pods, aggregation
over existing
ResourceClaim.status.allocation, or a dedicated scheduler-owned API resource — exact mechanism TBD.
Alternatives
CEL selector on each extractor
SharingAffinityExtractor could carry an optional Selector *DeviceSelector
— the same type used by DeviceRequest.selectors — restricting an extractor to the pool’s
devices whose attributes satisfy the expression. Extractors without a
selector would apply to every device. This enables heterogeneous
pools: a networking driver publishing both NICs and VFs in one pool
could bind subnet/pkey to NIC devices and parentNIC to VF
devices, dispatching on existing attributes such as device.attributes["type"].
Sharing affinity instead applies to every device in a pool; a driver whose devices need different affinity keys publishes them in separate pools.
Rejected because:
- No user story requires it: All three user stories (RDMA PKey alignment, FPGA bitstream sharing, single-subnet NIC sharing) are single-device-kind and need exactly one set of keys. Heterogeneous dispatch is a capability without a consumer.
- Pool splitting is a complete substitute: A driver chooses its own pool names, so publishing NICs and VFs as two pools needs no API support and no scheduler involvement. The selector does not add heterogeneous affinity, only heterogeneous affinity within a single pool — a packaging convenience rather than a missing capability.
- Uniqueness becomes a per-device property: Without selectors,
nameis unique across the whole list, so+listMapKey=namecatches duplicates within a slice declaratively and the cross-slice check is a string-set comparison. With selectors,namecan no longer be the list key — the same name may legitimately appear twice under disjoint selectors — so the declarative check is unavailable, and the authoritative check becomes “evaluate every selector against every device in the pool,” re-run whenever any slice in the pool changes. Both designs need allocator-side pool validation for the cross-slice case; the selector version makes that check O(extractors × devices) rather than O(extractors), and downgrades the diagnostic from “this pool is invalid” to “these devices in this pool are invalid.” See Key name uniqueness .
Per-device extractor reference
Instead of dispatching from the metadata slice, each Device could
carry a field naming which sharing-affinity entries apply to it (for
example sharingAffinityRefs []string matched against named
extractors) — explicit dispatch, no CEL evaluation, readable directly
from the device.
Rejected because:
- It solves a problem the design does not have: Like the selector, it exists only to serve heterogeneous pools, which are out of scope by decision.
- ResourceSlice size: The reference repeats on every device, in an object already capped at 128 devices and bounded by etcd limits, versus a single expression in the metadata slice.
- A dangling-reference surface: A device could name an extractor
that no metadata slice defines, which the API server cannot catch
since it validates each slice in isolation — exactly the problem
KEP-4815 documents for
DeviceCounterConsumptionreferences.
If that decision is ever revisited, the selector is the preferred mechanism: it reuses an existing type and adds no per-device bytes.
Well-known JSON schema inside OpaqueDeviceConfiguration
An earlier iteration of this KEP supplied claim-side affinity values via a
scheduler-recognized JSON schema embedded in OpaqueDeviceConfiguration —
i.e., a magic driver: resource.k8s.io opaque payload with
apiVersion: resource.k8s.io/v1alpha1, kind: StructuredParameters that the
scheduler would decode at Filter time.
Rejected because:
- Violates the contract of
Opaque:OpaqueDeviceConfigurationis, by name and design, opaque to the core API; only the driver is supposed to own its schema. A reserved global driver namespace decoded by the scheduler effectively turns part of the opaque payload into a core API surface without giving it API-server validation. - Weakly typed: A JSON-schema-inside-string approach is not strongly typed. API validation cannot enforce the schema, so most invariants shift to scheduler-side runtime decode/validation.
- Constrained evolution path: Schema versioning lives inside an opaque blob rather than in the Kubernetes API surface.
Typed Structured sibling on DeviceConfiguration
A subsequent iteration of this KEP added a typed Structured sibling on
DeviceConfiguration so the claim could carry affinity values in a
strongly-typed, API-validated map:
type DeviceConfiguration struct {
Opaque *OpaqueDeviceConfiguration
Structured *StructuredDeviceConfiguration // proposed
}
type StructuredDeviceConfiguration struct {
Requests []string
Parameters map[string]StructuredParameterValue
}
The scheduler would read Structured.Parameters directly to extract
affinity values, with no need to interpret opaque payloads.
Rejected because:
- New typed surface on a core API: Adds a new field to the
resource.k8s.ioAPI group that exists solely to serve scheduler-readable parameter extraction — a non-trivial API surface cost for a single consumer, when the underlying claim already carries the same data in driver-specific form viaOpaque. - Forces claim duplication: For most realistic claims the affinity keys
(subnet, pkey, vlan, …) are already present in the driver’s opaque config.
The
Structuredfield would require the claim author to duplicate those values in a second, scheduler-visible location, with no automatic enforcement that the two stay consistent. - Reviewer guidance (per @johnbelamaric, @pohly on PR #5987): The scheduler does not need a new typed claim-side field; it needs a way to extract affinity keys from the existing opaque config. The right mechanism for that extraction is driver-published CEL — which the driver already owns the schema for — placed on the ResourceSlice next to the device declarations.
Claim-side-only SharingAffinity (on DeviceRequest)
An alternative design adds a dedicated SharingAffinity field directly on
DeviceRequest within ResourceClaim, with no corresponding declaration on
the device or ResourceSlice. Sharing intent is expressed entirely from the
consumer side — the device is mute on whether or how it can be shared:
type DeviceRequest struct {
// ... existing fields ...
SharingAffinity *SharingAffinity
}
type SharingAffinity struct {
AffinityKey string
Value string
Strategy SharingStrategy
}
Rejected because:
- Wrong layer: Sharing affinity is a property of how a device can be
shared (a hardware-modal constraint declared by the driver), not a
property of an individual request. Putting it on
DeviceRequestwould imply the consumer chooses the strategy, when in reality the device (and its driver) dictates which keys must agree across consumers. - Single key only: The shape above implies one
AffinityKey/Valuepair per request. Multi-key affinity (e.g.,subnet+pkey+vlan) would require either repeating the field or introducing a list, both of which converge structurally to “a map of typed parameters” — which is exactly what the adopted CEL extraction produces, without requiring a new claim-side field at all.
Object Reference-based Affinity Matching
An alternative approach replaces inline affinity values with external
object references. Instead of extracting values from the claim’s opaque
config, the claim would reference a CRD (e.g., NetworkConfiguration) by
name, and the device would declare which object kinds constrain sharing.
Rejected because:
- Requires new fields on both ResourceClaim and Device (or ResourceSlice),
whereas the adopted approach adds only a single slice-level
sharingAffinityfield. - Requires external CRD definitions, adding operational burden for cluster administrators.
- Multi-dimensional affinity: A device may need affinity on multiple independent axes (e.g., subnet + VLAN). With object references, each axis would need its own CRD.
- Indirect object references raise authorization and lifecycle concerns (who owns the CRD instance? what happens when it is deleted while claims reference it?).
CEL on runtime scheduler lock state (rejected variant)
A related CEL-based design — considered before the design pivot — would have had the driver publish a CEL expression on the ResourceSlice that evaluates whether a claim is compatible with the device’s current lock state:
sharingAffinity:
lockExpression: >
device.affinityLock['subnet'] == '' ||
device.affinityLock['subnet'] == claim.AffinityValues['subnet']
Rejected because:
device.affinityLockis runtime scheduler state, not a static device attribute. Exposing it in CEL requires extending the evaluation context to include the scheduler’s in-memoryAllocatedState, which breaks the current model where CEL evaluates against the ResourceSlice snapshot.- CEL expressions are powerful but opaque to the scheduler — it cannot
extract which keys constrain sharing or what values to record in
AllocatedState. The scheduler would need to both evaluate the expression AND separately track lock state, duplicating logic. - A circular variant — where one claim’s eligibility depends on another claim’s allocation — produces non-deterministic results depending on evaluation order.
Future Enhancements
The following ideas are out of scope for alpha but are worth exploring in beta/GA based on real-world feedback:
Affinity-aware scoring (planned for Beta)
This KEP’s Filter phase is sufficient for correctness — an incompatible locked device is filtered out, and any remaining candidate produces a valid allocation. It is not, however, sufficient for packing, at either scope:
- Cross-node: stock Kubernetes scorers do not factor in
sharingAffinitywhen scoring nodes yet. Asubnet=Xclaim is just as likely to land on a node with a clean device as on a node that already has a compatibly-locked device. - Within-node: once a node is selected, the DRA structured-parameters allocator picks the first feasible device (first-fit). Among multiple feasible devices on the chosen node — say, one already locked to a compatible affinity value and one clean — there is no preference logic; the allocator may pick whichever appears first in the ResourceSlice.
The DynamicResources plugin already implements scoring for DRA
(shipped in K8s 1.35). Prioritized List (KEP-4816) and Extended
Resources (KEP-5004) already contribute their own additive terms to
computeScore. This is the canonical extension point for new DRA
features that need scoring.
KEP-5981 plans to add a sharing-affinity term to computeScore as a
beta deliverable, following the same per-feature additive pattern. The
broader “general-purpose DRA scoring” discussion continues in
kubernetes/enhancements#4970
,
but this KEP’s contribution does not block on a unified framework
or on #4970 producing a generic API. Doing per-feature scoring here is
consistent with how Prioritized List and Extended Resources already
shipped scoring, and avoids stranding sharing affinity behind a
multi-feature design effort.
The detailed score terms, weights, and tie-breakers will be designed and proposed as part of the beta graduation. Alpha intentionally does not commit to specific scoring shape so that beta has freedom to incorporate operational feedback from alpha.
Within-node device selection is addressed similarly by extending the allocator’s per-node selection logic so that, among feasible devices for a chosen sub-request, the allocator prefers a locked-compatible device over a clean one. This is a localized change to the structured-parameters allocator and is also a beta deliverable.
Alpha contract: correctness only. Packing of any kind (within-node or cross-node) is best-effort first-fit and is documented as a known limitation in Risks and Mitigations .
Priority-based Lock Preemption
Removed in PR review (2026-05): lock-breaking preemption is not a KEP-5981 deliverable. See Composition with DRA Preemption (KEP-5690) under Notes/Constraints/Caveats.
SharingStrategy (CanSetLock / NeverSetLock)
Alpha intentionally does not let claims control whether they may establish a new lock on a clean device. Any compatible claim can set the initial lock, and subsequent compatible claims can then reuse that locked device.
A future enhancement could add an explicit SharingStrategy on the claim side to control lock-setting behavior. Two candidate strategies are:
CanSetLock(default): The claim may land on a clean device and establish the lock. This matches the alpha behavior.NeverSetLock: The claim may only be allocated to a device that already has a matching lock established by another claim. This is useful for background or batch jobs that should never consume a clean device and potentially fragment capacity. Caveat:NeverSetLockis a follower-only strategy — it requires at least oneCanSetLockclaim to establish the lock first. If no device is locked to the requested value, aNeverSetLockpod will remain unschedulable indefinitely. Implementations should document this dependency clearly and consider surfacing a scheduling event when a pod is blocked waiting for a lock that no leader has established.
If introduced in beta or later, the scheduler would evaluate this policy before
capacity and key matching for unlocked devices. A claim with NeverSetLock
would reject an unlocked device immediately, then continue searching for an
already-locked compatible device.
This is deferred from alpha to keep the initial scope focused on the core problem: driver-declared sharing constraints plus scheduler-enforced lock tracking via CEL-extracted parameters.
Soft / Preferred Affinity Keys
The Alpha design enforces hard all-or-nothing matching across the pool’s extractors: every extracted key must agree with the device’s current lock, or the device is filtered out. Real-world hardware may have hierarchical constraints where some keys are strict sharing requirements (e.g., Subnet) and others are scheduling preferences (e.g., Traffic-Class or bandwidth profile).
A future enhancement could let the driver mark individual CEL
expressions as required vs preferred:
required(default): Mismatch → device filtered out; key contributes to the lock (current behavior).preferred: Mismatch → device passes Filter but is deprioritized in device selection; key does not contribute to the lock.
Infrastructure Needed
None