KEP-5981: DRA Sharing Affinity

Implementation History
ALPHA Implementable
Created 2026-03-30
Latest v1.38
Milestones
Alpha v1.38
Beta v1.39
Stable v1.41
Ownership
Owning SIG
SIG Scheduling
Participating SIGs
Primary Authors

KEP-5981: DRA Sharing Affinity

Release Signoff Checklist

Items marked with (R) are required prior to targeting to a milestone / release.

  • (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
  • (R) KEP approvers have approved the KEP status as implementable
  • (R) Design details are appropriately documented
  • (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
    • e2e Tests for all Beta API Operations (endpoints)
    • (R) Ensure GA e2e tests meet requirements for Conformance Tests
    • (R) Minimum Two Week Window for GA e2e tests to prove flake free
  • (R) Graduation criteria is in place
  • (R) Production readiness review completed
  • (R) Production readiness review approved
  • “Implementation History” section is up-to-date for milestone
  • User-facing documentation has been created in kubernetes/website , for publication to kubernetes.io
  • Supporting documentation—e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes

Summary

This KEP proposes an extension to Dynamic Resource Allocation (DRA) that allows the kube-scheduler to handle resources that are conditionally fungible.

KEP-5075 (Consumable Capacity) introduced the ability to track numerical capacity (e.g., 16 slots of a NIC) and share devices across multiple claims via allowMultipleAllocations. However, it assumes all claims are fungible—any claim can share the device with any other claim.

Real-world hardware is often modal (i.e., once partially allocated, it must operate in a single configuration mode for all of its current consumers): the device requires all subsequent consumers to share a specific configuration. For example:

  • Multi-pod NIC sharing: A network DRA driver shares a NIC across 16 pods, but all pods must belong to the same subnet. Once the first pod configures the NIC for Subnet A, the remaining 15 slots are restricted to Subnet A.
  • FPGA bitstream sharing: An FPGA can serve multiple inference pods, but all must use the same bitstream. Once bitstream-ml-v2 is loaded, other pods needing bitstream-crypto-v1 must use a different FPGA.

This KEP introduces a sharingAffinity field on ResourceSlice.spec that allows drivers to declare, via CEL expressions, how to extract the affinity keys that constrain sharing from the driver’s existing opaque device configuration. sharingAffinity is published as pool-level metadata in dedicated slices within the pool — a slice that carries sharingAffinity does not carry devices, and vice versa — mirroring the established sharedCounters (KEP-4815) pattern. A pool may publish one or more such slices; the scheduler takes the union of their extractors, so the schema declaration stays scoped to the pool even when the pool is chunked across many ResourceSlices.

On the claim side, there is no API change — workloads continue to author their driver’s opaque config exactly as they do today. When the scheduler is evaluating a candidate device, it looks up the pool’s metadata slice, runs every published CEL extractor against the opaque-config objects in scope for the request currently being filtered (per DeviceClaimConfiguration.Requests), and — if every extractor is satisfied — records the resulting key/value pairs in AllocatedState alongside consumed capacity. This enables the scheduler to gate remaining capacity on locked devices and safely reuse them for compatible claims when selected by the existing allocator. Alpha provides correctness only — affinity-aware preference (packing) is delivered in beta; see Goals .

In addition, if a device already has active allocations whose affinity cannot be reconstructed (for example, legacy claims created before the feature was enabled), the scheduler treats that device conservatively and does not place new sharingAffinity allocations on it until the device becomes clean.

sharingAffinity in this KEP refers specifically to compatibility for co-allocation on a shared device; it is distinct from pod affinity, anti-affinity, or topology-aware placement.

Motivation

As AI and HPC workloads move toward higher density, hardware partitioning (SR-IOV, GPU slicing, FPGA multi-tenancy) is becoming standard. These physical devices often have a “modal” constraint (see Summary for the definition and concrete examples).

Currently, the scheduler is unaware of this “lock.” It may schedule a Pod requiring a different configuration to the same device because it sees “available capacity.” In short: In these scenarios, Quantitative Sharing (how many slots?) fails without Qualitative Gating (what mode are those slots in?). This leads to:

  1. Allocation failures at the node level: The driver rejects incompatible binds at prepare time, after the scheduler has already committed
  2. High scheduling latency: The scheduler retries the same failing combination, thrashing between candidates
  3. Resource starvation: Without affinity awareness, same-subnet pods spread across multiple devices instead of consolidating—wasting capacity
  4. Complex driver workarounds: Drivers resort to placeholder patterns with race conditions and ResourceSlice churn (see Status Quo below)

The scheduler’s AllocatedState currently tracks consumed capacity but not the affinity values that determine sharing compatibility. This KEP closes that gap.

Status Quo: Driver-Side Placeholder Pattern

Without this KEP, drivers must use a “placeholder pattern” today:

  1. Publish devices with capacity: 1 initially
  2. Wait for first claim to determine affinity value
  3. Update ResourceSlice with actual capacity, writing the affinity value into the device’s attributes map
  4. Use CEL selector to match against that attributes entry

Problems:

  • Race condition: Second pod may go to different device before expansion
  • ResourceSlice churn: Constant updates as pods come and go
  • Driver complexity: State machine for expand/contract lifecycle

Goals

  • Enable the scheduler to gate remaining capacity on a device based on a required affinity key
  • Provide a mechanism for drivers to signal compatibility requirements for shared hardware via sharingAffinity on the ResourceSlice — without requiring drivers to change their existing opaque config schemas, and without requiring workload authors to learn a new claim-side API
  • Reduce fragmentation of cluster resources by enabling the scheduler to pack workloads with compatible sharing requirements onto already-locked devices (delivered in beta as a sharing-affinity term added to DynamicResources.computeScore — see Affinity-aware scoring (planned for Beta) ; alpha provides correctness only)
  • Track affinity values in AllocatedState so subsequent scheduling decisions respect the first claim’s lock-in
  • Maintain backward compatibility with devices that have no sharing affinity constraints, and with existing opaque-config workloads

Non-Goals

  • Defining hardware-specific affinity key names (these remain driver-defined)
  • Managing the physical lifecycle of the device configuration (this remains the driver’s responsibility)
  • Changing how capacity is tracked (that’s KEP-5075)
  • Supporting affinity across multiple devices. The lock is scoped to a single Device object in ResourceSlice.spec.devices[] — affinity is never shared across separate Device objects, even ones of the same type within the same pool, and certainly not across device types or pools.
  • Retrofitting affinity-aware sharing onto already-in-use devices when active claims do not expose reconstructable affinity values. In alpha, such devices are treated conservatively until they drain clean.
  • Guaranteeing lock-breaking preemption. This KEP does not introduce any preemption logic. Baseline DRA preemption is being introduced separately by KEP-5690 ; this KEP is designed to compose transparently with it (see Composition with DRA Preemption (KEP-5690) ) but does not depend on or deliver it.

Proposal

Add a sharingAffinity field to ResourceSlice.spec that publishes a list of driver-defined CEL extractors. Each entry names one affinity key and carries the CEL expression that produces its value, and every entry applies to every device in the pool. The scheduler evaluates an entry’s expression once for each opaque config in the claim, binding that config to a CEL variable named object — its JSON payload decoded into a nested map. (DRA already requires the payload to be valid JSON, so this KEP adds no new constraint.) The expression returns either a string value — the affinity key for that config — or the empty string "", meaning the extractor does not apply to it. A request satisfies an extractor when it produces a non-empty value across the request’s in-scope opaque configs, and the device is viable only if every extractor in the pool is satisfied. The non-empty returns from all extractors form the request’s effective affinity key map. If two opaque-config objects in the same request return different non-empty values for the same key, the claim is self-inconsistent and is treated as a user authoring error (see Extraction failure modes ).

Sharing affinity is published in dedicated metadata slices that are mutually exclusive with devices. Device-bearing slices in the same pool reference back to the metadata via the standard pool tuple (driver, pool.name, pool.generation). A pool may publish one or more metadata slices; the scheduler takes the union of sharingAffinity extractor entries across all metadata slices in the same complete pool (same generation, all resourceSliceCount members present, per the pool-completeness rule from KEP-4815 ). Before reading metadata, the scheduler invokes the existing pool-completeness check; devices in incomplete or invalid pools are treated as ineligible for sharing-affinity decisions until the pool becomes complete and clean. Conflict resolution rides on existing rules — incomplete or invalid pools are skipped entirely (KEP-4815’s fail-closed gate), and key names must be unique across all of the pool’s extractors (see Key name uniqueness ), so multi-metadata-slice publishing reduces to the same semantics as a single metadata slice.

# Metadata slice: carries sharingAffinity, no devices.
apiVersion: resource.k8s.io/v1
kind: ResourceSlice
metadata:
  name: networking-node-a-meta
spec:
  driver: networking.example.com
  nodeName: node-a
  pool:
    name: node-a
    generation: 7
    resourceSliceCount: 2
  sharingAffinity:
    - name: subnet
      expression: |
        object.apiVersion == "networking.example.com/v1"
          && object.kind == "NICConfig"
          ? object.subnetID : ""
    - name: pkey
      expression: |
        object.apiVersion == "networking.example.com/v1"
          && object.kind == "NICConfig"
          ? object.ibPKey : ""
---
# Device slice: carries devices, no sharingAffinity. Joined to the
# metadata slice by the (driver, pool.name, generation) tuple.
apiVersion: resource.k8s.io/v1
kind: ResourceSlice
metadata:
  name: networking-node-a-0
spec:
  driver: networking.example.com
  nodeName: node-a
  pool:
    name: node-a
    generation: 7
    resourceSliceCount: 2
  devices:
    - name: eth1
      allowMultipleAllocations: true
      capacity:
        networking.example.com/slots:
          value: "16"

Each CEL expression is responsible for returning the empty string when the opaque-config object is not one it applies to. How it decides is the driver’s choice; guarding on object.apiVersion and object.kind is the recommended convention, not a requirement. This contract gives drivers full control over which opaque-config schemas a given extractor applies to, and avoids forcing the scheduler to know anything about driver schema taxonomies. CEL runtime errors (for example, dereferencing a field that does not exist on the parsed config) are not equivalent to returning "" — they are ResourceSlice authoring errors, and the scheduler aborts allocation and fails scheduling for the Pod rather than skipping the device (see Extraction failure modes ). An apiVersion/kind guard is the simplest way to avoid this for configs that belong to other schemas.

If a driver supports multiple opaque-config schemas where the extracted fields live at the same path on each, one extractor can handle them all:

sharingAffinity:
  - name: subnet
    expression: |
      (object.apiVersion == "networking.example.com/v1" ||
       object.apiVersion == "networking.example.com/v1beta1")
        && object.kind == "NICConfig"
        ? object.subnetID : ""

All extractors published by a pool apply to every device in that pool; there is no per-device dispatch. A pool is therefore homogeneous with respect to sharing affinity — every device in it is gated on the same set of affinity keys. A driver whose devices need different affinity keys publishes them in separate pools.

A claim continues to author the driver’s opaque config exactly as it does today — no new claim-side API field is introduced:

config:
  - requests: ["nic"]
    opaque:
      driver: networking.example.com
      parameters:
        apiVersion: networking.example.com/v1
        kind: NICConfig
        vendor: acme
        subnetID: subnet-A
        ibPKey: "0x8001"
        qos: gold       # driver-private; not extracted, not seen by scheduler
        mtu: 9000       # driver-private
        vlanTag: 100    # driver-private

When the scheduler evaluates a multi-allocatable device backed by a pool that declares sharingAffinity:

  1. First request targeting this device: When evaluating a candidate device in this pool for a DeviceRequest, the scheduler looks up the pool’s metadata slice and, for each sharingAffinity entry, runs the entry’s CEL expression against every opaque config object that is in scope for the request (per DeviceClaimConfiguration.Requests). An extractor is satisfied when it produces a non-empty return across the request’s opaque configs. The device is a viable option for the request only if every extractor is satisfied; the union of their key/value contributions is the request’s effective affinity map and is recorded in AllocatedState alongside consumed capacity.
  2. Subsequent requests: The scheduler runs the same extraction for each new request targeting devices in the pool and compares the merged key set to AllocatedState’s LockedAffinity.
  3. Mismatch: If extracted keys do not match the recorded keys — as defined under Extraction failure modes — the device is not a viable option for that request and the scheduler tries another candidate device. Key sets always align, because Strict Gating admits a request only when every extractor is satisfied: a valid config must set every field the pool extracts.
  4. Match: If all extracted keys match and capacity is available, allocation proceeds.

Keys absent from the slice’s CEL expression set are not extracted, so driver-private config (e.g., qos, mtu, vlanTag above) flows through to NodePrepareResources untouched and never participates in lock evaluation. The driver remains the sole authority for those fields.

Alpha Design Decisions

1. Placement of extraction: ResourceSlice (driver-side, pool-level metadata slice)

sharingAffinity lives on ResourceSlice.spec and is mutually exclusive with devices. A pool publishes one or more metadata slices that carry sharingAffinity and zero devices; the pool’s device-bearing slices carry devices and no sharingAffinity. When multiple metadata slices exist in the same complete pool, the scheduler unions their extractor entries (per the multi-slice rule described above).

The driver is the natural owner: the extraction logic describes the driver’s own opaque config schema, which is uniform across all devices the driver publishes in this pool and evolves with driver versions, not with hardware instances. Pool-level placement avoids per-slice duplication, removes the drift risk of having to keep the extractor block in sync across many slices when a pool is chunked, and aligns the declaration site with what is being declared (a driver-schema property, not a device property). Per-device overrides can be added later if a driver ever needs them.

2. How affinity values reach the scheduler: CEL on opaque config

The scheduler does not interpret driver-private fields — it only evaluates the CEL expressions the driver has explicitly published as extraction logic. Each extractor is responsible for recognizing the schema(s) it applies to (typically by inspecting object.apiVersion / object.kind and returning "" when it does not).

// SharingAffinityExtractor declares one CEL expression that produces
// one sharing-affinity key from a driver's opaque device config.
// Every entry in a pool's `sharingAffinity` list applies to every
// device in the pool. See Design Details → API Enhancement for the
// canonical godoc.
type SharingAffinityExtractor struct {
    Name       string // affinity key name; unique within the pool
    Expression string // CEL over `object`; must return a string
}

Properties of this design:

  • No claim-side API change: workloads keep authoring the driver’s existing opaque config. No migration cost on the user side.
  • Minimal scheduler assumption: the scheduler’s only assumption is “run some CEL against the claim’s opaque configs and see what key/value pairs come back.” It never owns or interprets the driver’s schema.
  • Driver-private keys are invisible to the scheduler: anything the driver does not publish a CEL expression for is never extracted. The scheduler cannot inadvertently lock on hardware-configuration fields (qos, mtu, vlan) that drivers want to keep within their domain.

Extractor evaluation contract

An extractor’s CEL is evaluated against every opaque-config object in scope for the request being filtered. For each (extractor, opaque-config, key) triple:

  • A non-empty string return contributes that key → value to the request’s effective affinity map.
  • An empty string return contributes nothing for that key from that object. This is the idiom drivers use to express “this extractor does not handle this opaque-config schema.” Empty returns are not errors.

After evaluating every extractor against every opaque-config object, the scheduler checks whether each extractor is satisfied — it produces a non-empty return across the request’s opaque configs. The device is a viable option for the request only if every extractor in the pool is satisfied (Strict Gating).

Extraction failure modes

Three categories of outcomes can arise when evaluating a candidate device whose pool declares sharingAffinity for a request. They differ in who is responsible for the failure and what action the scheduler takes:

  1. Strict Gating outcome (normal, not an authoring error) — at least one extractor is not satisfied for this (request, candidate device) pair (the extreme case is “no keys produced at all”; a partial case is a pool whose subnet extractor resolved but whose pkey extractor did not). The device is not a viable option for the request — the safe default, since the driver declared that sharing requires scheduler-readable extraction across every published extractor. The request stays Pending; the pod may still schedule onto another candidate device in this or another pool, including pools that do not declare sharingAffinity.

  2. ResourceClaim authoring errors → UnschedulableAndUnresolvable — the claim’s own opaque configs are internally inconsistent; no slice the driver might publish can fix this.

    • Self-inconsistent extraction for a request: two (extractor, opaque-config) pairs in scope for the request produce the same key with different non-empty values. The pod is marked UnschedulableAndUnresolvable; the failure scope is the ResourceClaim, not a single device, because every device backed by a sharingAffinity pool that runs the same extractor against the same in-scope opaque configs will see the same contradiction. An Event is emitted on the pod for diagnosability. The user must fix the claim’s opaque configs before rescheduling.
  3. ResourceSlice authoring errors → abort allocation, fail scheduling for the Pod — the slice’s CEL is the problem; the driver must republish the slice to fix it.

    • Duplicate key name across extractors: two extractors in the pool declare the same key name. The lock state would be ambiguous about which extractor’s value to store. Within a single slice this is rejected at admission; across metadata slices in the same pool it is caught by pool validation. See Key name uniqueness .
    • CEL evaluation error or non-string return: a CEL expression fails to evaluate or returns a non-string. The scheduler does not silently lock the device to an empty key set.
    • Cost-budget exhaustion: CEL evaluation exceeds the per-eval cost budget.
    • Missing field on the parsed object: a CEL expression dereferences a field absent from object (e.g., object.subnet on an object with no subnet), raising a CEL runtime error. Indicates the extractor’s CEL did not defensively guard the access (e.g., should be has(object.subnet) ? object.subnet : ""). Typical causes: typo (wrong field name) or schema evolution where the driver renamed or removed a field across versions. Inapplicable-object cases (driver-tuning config, config for a different driver kind) should be expressed via has() guards returning "", not surfaced as errors.

    In every ResourceSlice-author-error case the scheduler aborts allocation and fails scheduling for the Pod instead of skipping the device and trying the next candidate. Skipping would emit an Event per candidate device and degrade silently into “no device matched”. The trade-off is that one malformed expression blocks the whole pool; admission-time validation of CEL syntax, cost, and length catches most errors before a slice is accepted.

    The canonical extractor idiom — guard with apiVersion / kind (or has(...)) and return "" from the non-applicable branch — distinguishes “this object intentionally does not contribute” ("") from “this object should have contributed but the schema does not match” (runtime error). Only the latter is treated as a slice authoring error.

Alpha scope

Alpha fully resolves the design around driver-side slice-level extraction described above. Claims do not control lock-setting behavior: any compatible claim may establish the initial lock on a clean device. Claim-side lock-setting policy (for example, CanSetLock/NeverSetLock) is deferred to Future Enhancements .

Alpha standardizes driver-declared CEL extraction, a single feature gate that governs both the slice-level field and the scheduler logic, and correct lock enforcement on already-locked devices — but intentionally stops short of affinity-aware scoring (planned as a beta-scope contribution to DynamicResources.computeScore, see Affinity-aware scoring (planned for Beta) ).

Alpha limitations

Alpha enforces lock compatibility but does not preempt incompatible lock-holders. A higher-priority Pod requiring a different affinity value than the current lock will remain unschedulable on that device until the lock-holder exits or a compatible alternative appears. This is consistent with baseline DRA, which has no preemption support in the absence of KEP-5690 . See Composition with DRA Preemption (KEP-5690) for how this KEP is structured to benefit transparently when KEP-5690 is present.

Alpha also restricts CEL return values to string. List-valued return types (a claim accepting any of several values for one key) are a recognized future extension but are out of scope for alpha — all motivating use cases (subnet IDs, PKeys, bitstream identifiers, NUMA tags) are single-valued.

A related extensibility note (raised by @pohly in PR review): a scalar string return type forecloses carrying any metadata alongside the extracted value. A future revision could return a structured object (for example a CEL message/struct with fields such as value, qualifiers, or per-key options) so that the contract can grow without another API break. Alpha deliberately stays on string because every motivating use case is a single opaque identifier and a struct return adds CEL-side type plumbing and scheduler-side parsing that is not justified by current requirements; the option is preserved by treating the return type as a versioned part of the extractor contract rather than a free parameter. See Future Enhancements for the extension path.

User Stories

Story 1: RDMA Partition Key Alignment

A user runs a distributed training job where every Pod must share the same RDMA Partition Key (PKey) to communicate. The NIC supports 16 VFs. The driver publishes sharingAffinity on the slice with a CEL expression pulling the PKey out of its existing opaque config (e.g., pkey: "object.ibPKey"). The scheduler only co-allocates Pods whose claimed PKey matches the NIC’s current lock (or selects an unlocked NIC and establishes the lock from the first claim).

  • Pod A (pkey-0x8001) is allocated to mlx5_0 → mlx5_0 is now locked to pkey-0x8001
  • Pod B (pkey-0x8001) arrives → matches affinity, is eligible to share mlx5_0
  • Pod C (pkey-0x8002) arrives → affinity mismatch on mlx5_0; mlx5_0 is not a viable option for Pod C; Pod C is allocated to mlx5_1 instead

Story 2: FPGA Bitstream Sharing

An inference service uses FPGAs to accelerate a specific model. Loading a bitstream takes several seconds. The driver publishes a CEL expression on the slice that extracts the bitstream identifier from its opaque config (e.g., bitstream: "object.bitstreamID"). The scheduler only co-allocates Pods that request a compatible bitstream onto an already-locked FPGA; affinity-aware preference for FPGAs that already have the bitstream loaded (over fresh ones) is delivered in beta.

  • Pod A (bitstream-ml-v2) is allocated an FPGA → FPGA locks to bitstream-ml-v2
  • Pod B (bitstream-ml-v2) arrives → eligible to share the same FPGA
  • Pod C (bitstream-crypto-v1) arrives → the locked FPGA is not a viable option for Pod C; uses a different FPGA or waits

Story 3: Single-subnet NIC Sharing

A network DRA driver advertises NICs that can be shared across up to 16 pods, but only if pods belong to the same subnet. The driver publishes a CEL expression on the slice that extracts the subnet from its opaque config (e.g., subnet: "object.subnetID").

  • Pod A (subnet-X) is allocated to eth1 → eth1 is now locked to subnet-X
  • Pod B (subnet-X) arrives → matches affinity, is eligible to share eth1
  • Pod C (subnet-Y) arrives → affinity mismatch on eth1; eth1 is not a viable option for Pod C; Pod C is allocated to eth2 instead

Notes/Constraints/Caveats

  • Affinity is set by the first compatible claim on a clean device: Once a device is allocated with an affinity value, that value is locked until all claims release the device.
  • Extractors: The pool’s sharingAffinity lists one or more extractors, each applying to every device in the pool. Evaluation semantics (Strict Gating across all of the pool’s extractors) are detailed in Filter Phase and Device Selection and Key name uniqueness .
  • Driver-private keys are invisible to the scheduler: Keys not present in the slice’s sharingAffinity list (e.g., qos, mtu, vlan) are not extracted and do not participate in lock evaluation. They flow through opaquely to the driver at NodePrepareResources as today. This is the mechanism by which the scheduler stays out of driver-private config.
  • String-only matching in alpha: CEL expressions must return string values. Non-string returns (numbers, booleans, lists, objects) are treated as extraction failures. Workloads use the normalized string form of their identifiers (subnet IDs, PKey hex strings, bitstream names, FQDNs).
  • CEL evaluation errors: A CEL expression that returns a non-string, errors, or exceeds the per-evaluation cost budget is a ResourceSlice authoring error: the scheduler aborts allocation and fails scheduling for the Pod (see Extraction failure modes ). The scheduler never silently establishes an empty lock when extraction fails.
  • Devices opting out of SharingAffinity extraction: Devices in slices without sharingAffinity behave as before — any claim can share them regardless of opaque config content.
  • Legacy allocations with unknown affinity are conservative in alpha: If a device has active allocations for which the scheduler cannot reconstruct the required affinity values (for example, claims created before the feature was enabled, or whose opaque configs produce no non-empty extractor returns under the currently-published CEL), that device is treated as having unknown affinity state and is filtered out for new sharing-affinity scheduling until it becomes fully clean.

Composition with DRA Preemption (KEP-5690)

This KEP does not deliver preemption. Lock-breaking preemption is not a KEP-5981 deliverable in any milestone.

The design is structured so lock-breaking falls out transparently when KEP-5690 (DRA Preemption) is present in the cluster: affinity locks live in the same dynamicresources.stateData structure that KEP-5690 mutates via AddPod/RemovePod. The only implementation requirement on the KEP-5981 side is that AddPod/RemovePod correctly mutate affinity-lock state alongside capacity — which is already required for KEP-5690 compatibility.

Known limitation: affinity-blind reprieve ordering. SelectVictimsOnNode’s reprieve order is (priority, PDB, runtime) and does not consider which Device a Pod’s claim is allocated to. When a node has multiple candidate devices each holding a lock to a different value, with asymmetric lock-holder counts, the algorithm may converge on a victim set on the larger-lock-count device when sacrificing the smaller one would have sufficed. Always correct, sometimes over-evicts. A DRA-aware reprieve-ordering hook in DefaultPreemption would resolve this cleanly but is its own scheduler-framework enhancement, out of scope for this KEP.

Reclaiming an affinity-locked device is addressed by workload-aware preemption (cluster-wide victim identification), which is out of scope for KEP-5690’s initial integration and left to future enhancements.

If KEP-5690 is not present or is disabled, this KEP provides no preemption capability — see Preemption Cannot Break Affinity Locks .

Handling Legacy Claims with Unreconstructable Affinity

Device StateNew ClaimResult
5 legacy claims, affinity unknownClaim whose CEL extraction yields subnet: AFiltered out. Existing allocations have unknown affinity, so no new sharing-affinity lock may be established yet.
5 legacy claims, affinity unknownClaim whose extraction yields no non-empty keysFiltered out. Missing required scheduler-readable affinity information.
Legacy claims drained; device now cleanClaim whose CEL extraction yields subnet: ALock set to subnet: A; device now locked.
Device locked to subnet: AClaim whose CEL extraction yields subnet: AAllowed (values match).
Device locked to subnet: AClaim whose CEL extraction yields subnet: BFiltered out (mismatch with lock).
All claims released—Device fully clean and eligible to establish a new lock.

Legacy claims continue to run and are not evicted. However, until all unknown allocations on a sharing-affinity device are released, the scheduler does not assume it knows the device’s effective modal state.

Changing sharingAffinity on a Slice with Active Allocations

In alpha, mutating a slice’s sharingAffinity (adding, removing, or changing extractor entries or CEL expressions) while bound claims reference devices in that slice is not supported. On the next reconciliation or scheduler restart, the scheduler re-runs the new CEL against each bound claim’s opaque config. If any bound claim no longer yields the same key set as the current LockedAffinity (because a key was added, removed, renamed, or the CEL now returns a different value), the device is marked AffinityStates[deviceID].Status = AffinityStatusUnreconstructable and filtered out for new sharing-affinity scheduling until all such claims drain.

This is the same conservative-fallback behavior used for legacy claims and is the deliberate alpha trade-off: the scheduler refuses to silently downgrade its safety guarantee in the face of an asymmetric extraction change. Drivers that need to evolve sharingAffinity for in-use slices should drain affected devices before publishing the change, or stage schema evolution by adding a new extractor entry (e.g., one whose CEL guards on a new apiVersion) and migrating workloads to author against the new schema before retiring the old extractor.

Driver responsibility: drivers should avoid hot-swapping sharingAffinity entries on slices with active allocations. When extraction changes are unavoidable (e.g., a hardware capability evolves), drivers should expect the affected devices to be ineligible for new affinity-aware scheduling until they drain clean, and should plan rollouts accordingly (for example, by cordoning the device or rolling out the extraction change as part of a node reimage).

Compatibility Matrix

To clarify the interaction between requests and devices, the following matrix outlines how the scheduler and driver evaluate candidates based on whether the device’s pool declares a sharingAffinity Affinity Extractor list (AE) and whether the request’s opaque configs satisfy every extractor in that pool’s sharingAffinity list (GS — under Strict Gating every extractor in the pool must be satisfied):

ScenarioPool AERequest GSScheduler OutcomeDriver Outcome
Standard Feature UseYesYesMatch enforced. Extracted values match lock + capacity available → request scheduled.Validates hardware mode matches claim config at NodePrepareResources. Rejects if stale or inconsistent.
Strict GatingYesNoNot viable. Device is not a viable option for the request — at least one extractor was not satisfied by the request’s opaque configs (anything from “no keys at all” to “one key missing”). Request stays Pending; may schedule elsewhere.N/A — request never reaches the driver for this device.
Legacy Device TransitionYes (newly added)YesNot viable while legacy claims are active (Status: Unreconstructable). Allowed once device drains clean.Validates as normal once request reaches the driver. During transition, driver continues serving legacy claims.
Permissive SharingNoYesAllowed. Pool has no sharingAffinity; opaque config is not evaluated for affinity. Standard capacity matching applies.Must enforce hardware compatibility independently. Scheduler provides no affinity gating for this device.
Legacy/BasicNoNoAllowed. Standard DRA capacity and attribute matching.Must enforce hardware compatibility independently. This is the pre-KEP-5981 behavior.
Gate Disabled in SchedulerYesN/ANot viable. The scheduler cannot evaluate affinity, so every device in the pool is skipped rather than allocated with enforcement silently off (see Feature Gates ).N/A — request never reaches the driver for this device.

The top rows show the scheduler as the primary enforcer with the driver as a backstop. The bottom rows show the driver as the sole enforcer with the scheduler being permissive. The transition row shows the scheduler being conservative (filtering) while the driver continues serving existing workloads. The last row is the fail-closed case: a pool asked for enforcement the scheduler cannot provide, so its devices are withheld rather than handed out unenforced.

Risks and Mitigations

Fragmentation (Poisoning)

Risk: A claim with a rare or unique affinity value can lock a high-capacity device to that value, stranding the device’s remaining capacity against peer claims (any priority) that don’t share the value. Pure capacity-stranding, not a priority problem — even claims of equal or lower priority cannot use the locked device. This section covers the peer-claim case only; the priority-aware case is Preemption Cannot Break Affinity Locks .

Mitigation (Alpha): None beyond Filter correctness. The DRA allocator currently uses a first-fit algorithm with no affinity-aware preference, so the scheduler does not actively pack compatible claims onto already-locked devices. Where domain-specific validation is feasible, cluster administrators can use DeviceClass CEL selectors to restrict which affinity values are accepted (e.g., constraining subnet IDs to a known set) — this is an admin-side guardrail against rare or arbitrary values poisoning devices. Drivers that ship a typed opaque-config CRD can additionally enforce the same constraints upstream via a validating admission webhook on the config object, rejecting claims carrying out-of-policy values before they ever reach the scheduler. Affinity-aware preference (within-node) and a Score contribution (cross-node) are planned as a beta-scope addition to DynamicResources.computeScore, following the same per-feature additive pattern that Prioritized List (KEP-4816, shipped in 1.35) and Extended Resources (KEP-5004) already use; see Affinity-aware scoring (planned for Beta) . Until that lands, fragmentation mitigation is best-effort and depends on the existing first-fit ordering of devices in ResourceSlices. General-purpose DRA scoring discussion continues in kubernetes/enhancements#4970 , but KEP-5981’s contribution does not block on a unified framework.

Preemption Cannot Break Affinity Locks

Risk: Standard Kubernetes preemption triggers on resource shortage, not on affinity-lock mismatch. A higher-priority Pod requiring a different affinity value than the current lock cannot evict the lock-holder under standard preemption — capacity is technically free, so preemption is never triggered. This is the same gap that affects all of DRA today, not specific to this KEP.

Mitigation: This KEP does not deliver preemption. The mitigation within this KEP’s scope is scoring/packing — see Fragmentation (Poisoning) and the Beta scoring contribution in Affinity-aware scoring (planned for Beta) . Lock-breaking is a property that emerges when DRA preemption is present in the cluster; see Composition with DRA Preemption (KEP-5690) .

Design Details

API Enhancement

ResourceSlice Spec

type ResourceSliceSpec struct {
    // ... existing fields (Driver, NodeName, Pool, Devices, etc.) ...

    // SharingAffinity declares the affinity keys that gate sharing
    // for every device in this pool, and the CEL expressions that
    // extract them from a claim's opaque config. A ResourceSlice may
    // set either Devices or SharingAffinity, not both; a pool carries
    // its sharing affinity in dedicated metadata slices alongside its
    // device-bearing slices. At most 8 affinity keys may be declared
    // per slice. See SharingAffinityExtractor for semantics.
    //
    // +optional
    // +listType=map
    // +listMapKey=name
    // +k8s:maxItems=8
    SharingAffinity []SharingAffinityExtractor
}

// SharingAffinityExtractor declares one sharing-affinity key and the
// CEL expression that produces its value from a driver's opaque
// device config. Every entry in a pool's `sharingAffinity` list
// applies to every device in the pool, and a request must satisfy
// every entry for a device in that pool to be viable.
type SharingAffinityExtractor struct {
    // Name is the sharing-affinity key name (for example "subnet").
    // Must be unique within the pool.
    //
    // +required
    Name string

    // Expression is the CEL expression evaluated for this key, over
    // the claim's opaque-config object bound as `object`. It must
    // return a string. An empty-string return means "no contribution
    // for this key from this object" — the idiom an extractor uses to
    // ignore opaque-config objects whose apiVersion / kind it does
    // not handle.
    //
    // The length of the expression must be smaller or equal to 10 Ki.
    // The cost of evaluating it is also limited based on the
    // estimated number of logical steps; the combined cost of all
    // extractors in a pool is capped by a shared CEL cost budget.
    //
    // +required
    Expression string
}

const SharingAffinityMaxEntries = 8

// Expression length reuses the existing DRA limit,
// CELSelectorExpressionMaxLength (10 Ki).

// SharingAffinityCELMaxCost is the maximum combined execution cost
// allowed for all sharing-affinity expressions in a single pool,
// mirroring DeviceClaimDerivedAttributeCELMaxCost. Tying the
// collective cost of a pool's extractors to the cost allowed for one
// CEL selector keeps worst-case Filter latency comparable to existing
// device selection.
const SharingAffinityCELMaxCost = 1000000

Scheduler Enhancement

Source of Truth for Affinity Locks

The scheduler derives affinity locks solely from CEL extraction over active claims’ opaque configs — not from device attributes on the ResourceSlice. The driver is NOT required to write locked affinity values back to the ResourceSlice.

  • The ResourceSlice declares how to extract sharing-affinity key/value pairs (sharingAffinity).
  • The claims declare what values they need (via their normal opaque OpaqueDeviceConfiguration.parameters blob — driver’s existing schema).
  • The scheduler combines these by running CEL at Filter time and maintains the lock in AllocatedState.

This avoids two sources of truth that could diverge, eliminates ResourceSlice churn (no update every time a lock is set/cleared), and keeps driver implementation simple. Drivers MAY optionally publish current locked values as regular device attributes for observability (e.g., visible via kubectl), but the scheduler does not depend on them.

When the last claim on a device is released, the scheduler clears the lock. The scheduler’s notion of “clean” is allocation-clean, not hardware-ready — a device is considered clean once no allocated claim references it, regardless of whether the driver has finished in-flight hardware reconfiguration on the node. The driver remains responsible for device lifecycle: tearing down the old configuration (via NodeUnprepareResources) and reconfiguring for new claims (via NodePrepareResources). Driver-level prepare/unprepare sequencing is the authoritative guard against reuse before reconfiguration completes; drivers that need stronger guarantees should hold their own per-device readiness state and reject prepare calls until reconfiguration is complete.

Safety Model and Responsibility Split

This feature intentionally keeps placement knowledge and hardware enforcement separate:

  • Scheduler guarantee: when it has successfully extracted affinity keys (via the slice’s sharingAffinity CEL) for all active allocations on a device, it will not intentionally co-place claims with incompatible affinity values on that device.
  • Conservative fallback: if the scheduler cannot reconstruct the effective affinity state of a device (for example, due to legacy claims, opaque configs that no longer produce any non-empty extraction under the current sharingAffinity entries, or CEL evaluation failures), it treats that device as unknown and filters it out for new sharing-affinity placements until the device becomes clean.
  • Driver guarantee: the driver remains the final authority for programming and validating the actual hardware mode during NodePrepareResources.
  • Failure handling: stale scheduler state or races may still cause prepare-time rejection, and that rejection remains the final safety backstop.
Cache Extension: Effective Device State

To prevent race conditions during high-volume scheduling, the scheduler maintains affinity locks in its internal cache rather than relying on API server round-trips. This is consistent with how DRA already handles capacity tracking via inFlightAllocations.

The scheduler’s AllocatedState is extended to track affinity values alongside consumed capacity:

// AffinityStatus is scheduler-internal cache state, never serialized.
// The zero value is AffinityStatusClean, so a device absent from
// AffinityStates is correctly Clean.
type AffinityStatus int

const (
    // AffinityStatusClean: no active claims on the device.
    // LockedAffinity is nil.
    AffinityStatusClean AffinityStatus = iota

    // AffinityStatusLocked: at least one claim is active and the
    // scheduler has reconstructed the device's affinity values.
    // LockedAffinity holds the current lock.
    AffinityStatusLocked

    // AffinityStatusUnreconstructable: at least one active claim's
    // affinity values cannot be reconstructed (CEL extraction
    // failed, did not produce a non-empty value for some declared
    // key, or claim predates the feature). The device is filtered
    // for new sharing-affinity placements until it becomes Clean.
    // LockedAffinity is nil.
    AffinityStatusUnreconstructable
)

type AffinityState struct {
    // Status is the device's current affinity state. See the
    // AffinityStatus constants for the meaning of each value and
    // when LockedAffinity is populated.
    Status AffinityStatus

    // LockedAffinity holds the device's affinity lock when
    // Status == AffinityStatusLocked. Nil for Clean and
    // Unreconstructable.
    LockedAffinity map[string]string
}

type AllocatedState struct {
    AllocatedDevices         sets.Set[DeviceID]
    AllocatedSharedDeviceIDs sets.Set[SharedDeviceID]
    AggregatedCapacity       ConsumedCapacityCollection

    AffinityStates map[DeviceID]AffinityState
}
Filter Phase and Device Selection

Filter phase: For a given node, the scheduler evaluates each device. A device whose pool declares sharingAffinity (looked up via the pool’s metadata slice) is a candidate ONLY if:

  1. It has sufficient consumable capacity (KEP-5075).
  2. The device’s AffinityStates[deviceID].Status is not AffinityStatusUnreconstructable.
  3. For each extractor in the pool’s sharingAffinity list, the scheduler evaluates the extractor’s CEL expression against every opaque-config object in scope for this request (per-request scoping via DeviceClaimConfiguration.Requests). Every CEL expression must evaluate successfully and return a string value. CEL evaluation errors, non-string returns, missing fields, or cost-budget exhaustion are ResourceSlice authoring errors: the scheduler aborts allocation and fails scheduling for the Pod so the problem surfaces immediately, rather than skipping the device (see Extraction failure modes ).
  4. Every extractor is satisfied — each produces a non-empty return across the request’s opaque configs (Strict Gating). The contributions from all extractors are merged into the request’s effective affinity map. Two contributions producing the same key with different non-empty values are a ResourceClaim authoring error (self-inconsistent claim); the pod is marked UnschedulableAndUnresolvable. Two extractors in the same pool declaring the same key name are rejected before this point — at admission within a slice, and by pool validation across slices (see Key name uniqueness ).
  5. The device’s AffinityStates[deviceID].LockedAffinity is either empty (unlocked) OR matches the request’s effective affinity map exactly for ALL extracted keys.

If a device has AffinityStates[deviceID].Status == AffinityStatusUnreconstructable, or if not every extractor is satisfied by the request, or extraction fails (any reason in #3), the device is not a viable option for the request. This is the safe default: the driver declared that sharing requires scheduler-readable extraction, and a scheduler that cannot reconstruct the current or requested affinity state cannot evaluate placement safely. Requests that do not need sharing-constrained devices should target devices in slices without sharingAffinity.

Device selection within a node: Among feasible devices on a chosen node, alpha does not introduce affinity-aware preference. Device selection continues to use the existing structured-parameters allocator (first-fit). This means that on a node with both a compatibly-locked device with capacity and a clean device, the allocator may pick whichever appears first in the ResourceSlice rather than preferring the locked one.

Affinity-aware preference (within a node) and node-level scoring (across nodes) are planned for beta — see Affinity-aware scoring (planned for Beta) .

Key name uniqueness

Affinity key names must be unique across all extractors in a pool — otherwise the lock state on a device would be ambiguous about which extractor’s value to store. Because every extractor applies to every device in the pool, uniqueness is a property of the key names alone: a set comparison, with no need to consider which devices an extractor applies to. Two enforcement layers cover it:

  • Admission-time (exact, within a slice): because sharingAffinity is a flat list keyed by name (+listType=map +listMapKey=name), duplicate key names within a slice are rejected by generic list validation — no sharing-affinity-specific validation code is required.
  • Pool validation (across slices): a pool may publish several metadata slices, and the API server validates each slice in isolation, so a collision between extractors in different slices of the same pool cannot be caught at admission. This is the same limitation KEP-4815 has for counter-set references, and it is handled the same way: the allocator validates complete pools and treats a duplicate key name as an invalid pool, failing closed for every device in it rather than silently picking one extractor’s value.

Uniqueness is therefore a static property of the pool, checked before any device is evaluated — not a per-device runtime condition.

Expression limits

Each extractor’s CEL expression is bounded the same way DRA already bounds its other driver-published CEL:

  • Length: at most CELSelectorExpressionMaxLength (10 Ki) per expression — the limit already applied to CELDeviceSelector and DeviceDerivedAttribute expressions.
  • Cost: the estimated evaluation cost of a pool’s extractors is capped by a shared budget, SharingAffinityCELMaxCost (1,000,000, roughly 0.1 s), mirroring DeviceClaimDerivedAttributeCELMaxCost. A shared budget rather than a per-expression one keeps the cost of a pool’s whole extractor set comparable to a single CEL selector, which matters because Filter runs every extractor against every in-scope opaque config for every candidate device.

Both limits are validated at admission per slice, and the shared budget is re-checked across a pool’s metadata slices by the same pool validation that enforces key-name uniqueness. Following the existing DRA convention, validation happens only when an expression is set or changed, so changing the limits in a future release leaves already published expressions valid; the scheduler additionally enforces the cost limit at runtime.

Reserve Phase: Tentative Locking

Once a node/device is selected, the Reserve plugin establishes a “tentative lock” in the scheduler cache before the Binding phase. Reserve reuses the key map Filter already extracted — no second CEL pass — and atomically:

  • If the device is still unlocked: record the claim’s keys as the device’s LockedAffinity and proceed.
  • If the device became locked since Filter (another pod won the race) or was already locked: re-check the cached keys against the current LockedAffinity. On match, co-allocate. On mismatch, fail Reserve and let the scheduler retry with the next candidate.

The tentative lock is immediately visible to subsequent scheduling cycles: if Pod-B is evaluated milliseconds after Pod-A’s Reserve (before Pod-A’s bind reaches the API server), Pod-B’s Filter phase sees Pod-A’s tentative lock and either joins it or skips the device.

If scheduling later fails for the pod (Unreserve), the tentative lock is removed unless another already-bound claim is still co-located on the device.

Scheduler Restart: State Reconstruction

On scheduler restart, the in-memory AffinityStates map is empty and must be rebuilt from already-cached state (bound ResourceClaims and their ResourceSlices) before the first scheduling cycle. The behavior contract is:

  • Recovery: for each bound claim on a device whose pool declares sharingAffinity (looked up via the pool’s metadata slice), the scheduler re-derives the claim’s key map and records it as the device’s LockedAffinity. No new API calls are required.
  • Conservative fallback on ambiguity: if extraction fails, yields no keys, or yields inconsistent values across claims sharing the device (which by construction should not happen, but may arise from historical bugs, manual etcd edits, or version skew), the device is marked AffinityStatusUnreconstructable and excluded from new sharing-affinity placements until all claims on it drain. The scheduler never infers a lock from ambiguous data.
  • No new persistence: reconstruction uses the same informer-cached data the scheduler already consumes; no migration, no new API surface.

Feature Gates

This KEP introduces one feature gate:

  • DRASharingAffinity (alpha): adds the sharingAffinity field on ResourceSlice.spec, the AllocatedState.AffinityStates cache dimension, and the scheduler Filter / Reserve logic that runs CEL extraction over claim opaque configs and matches the result against the lock. The gate must be enabled on both kube-apiserver and kube-scheduler for the feature to function. Asymmetric enablement fails closed rather than silently disabling enforcement:

    • apiserver on, scheduler off: drivers can write sharingAffinity and the apiserver persists it, but the scheduler cannot evaluate it. Rather than allocating those devices with affinity silently unenforced, the scheduler treats every device in a pool that declares sharingAffinity as unallocatable. This mirrors how DRAPartitionableDevices skips slices using sharedCounters and devices using consumesCounters when that gate is off.

    • apiserver off, scheduler on: per standard alpha-field handling, the apiserver strips sharingAffinity from new writes. Slices persisted during a prior enabled period are still served on read, so the scheduler continues to honor existing locks. New attempts to declare affinity do not take effect, but they are not silent: the ResourceSlice controller helper reports a DroppedFieldsError naming DRASharingAffinity, the same mechanism that already surfaces DRAPartitionableDevices and DRADeviceTaints.

    Failing closed costs capacity — a scheduler rolled back while pools still declare affinity cannot place new requests on those devices — but ignoring the field is worse: it lets the scheduler co-locate claims the driver declared incompatible, which the driver backstop then rejects at NodePrepareResources, producing a schedule/reject loop. Operators should have drivers stop publishing sharingAffinity before disabling the gate in the scheduler.

    The expected rollout order mirrors other DRA gates: enable on the apiserver first, then the scheduler; disable in reverse.

Because the claim side carries no new API surface (workloads keep their existing opaque configs), no separate API-only gate is needed.

Examples

ResourceSlice with Sharing Affinity

apiVersion: resource.k8s.io/v1
kind: ResourceSlice
metadata:
  name: node1-nics
spec:
  driver: networking.example.com
  nodeName: node1
  sharingAffinity:
    - name: subnet
      expression: |
        object.apiVersion == "networking.example.com/v1"
          && object.kind == "NICConfig"
          ? object.subnetID : ""
  devices:
    - name: eth1
      allowMultipleAllocations: true
      attributes:
        networking.example.com/type:
          string: "sriov-vf"
      capacity:
        networking.example.com/slots:
          value: "16"
    - name: eth2
      allowMultipleAllocations: true
      attributes:
        networking.example.com/type:
          string: "sriov-vf"
      capacity:
        networking.example.com/slots:
          value: "16"

The sharingAffinity block declares a single key (subnet) to be pulled from any opaque-config object that identifies itself as networking.example.com/v1 / NICConfig. The CEL guard on object.apiVersion and object.kind ensures the expression returns "" (no contribution) when applied to an opaque-config object from a different schema, rather than failing with a missing-field runtime error. See the Proposal section for the canonical guard idiom and the multi-version variant.

ResourceClaim (status quo opaque config)

apiVersion: resource.k8s.io/v1
kind: ResourceClaim
metadata:
  name: pod-a-nic
spec:
  devices:
    requests:
      - name: nic
        exactly:
          deviceClassName: shared-nic
    config:
      - requests: ["nic"]
        opaque:
          driver: networking.example.com
          parameters:
            apiVersion: networking.example.com/v1
            kind: NICConfig
            subnetID: subnet-X        # extracted as `subnet` by the slice's CEL
            vlanId: 100               # driver-private; not extracted
            mtu: 9000                 # driver-private; not extracted

Note: The claim shape is unchanged from pre-KEP-5981 DRA — the workload author writes the driver’s existing opaque config. The scheduler runs the slice-declared CEL expression against the parsed parameters — the apiVersion/kind guard matches, so the expression returns subnet-X and the scheduler records subnet: subnet-X for affinity lock evaluation. The driver-private fields (vlanId, mtu) flow through to NodePrepareResources untouched.

Multi-key Sharing Affinity Example

This example illustrates the alpha semantics when a slice extracts multiple keys.

A driver advertises shared RDMA-capable NICs where both subnet and PKey must match for pods to share the same device:

apiVersion: resource.k8s.io/v1
kind: ResourceSlice
spec:
  driver: networking.example.com
  nodeName: node1
  sharingAffinity:
    - name: subnet
      expression: |
        object.apiVersion == "networking.example.com/v1"
          && object.kind == "NICConfig"
          ? object.subnetID : ""
    - name: pkey
      expression: |
        object.apiVersion == "networking.example.com/v1"
          && object.kind == "NICConfig"
          ? object.ibPKey : ""
  devices:
    - name: mlx5_0
      allowMultipleAllocations: true
      capacity:
        networking.example.com/slots:
          value: "16"

A matching claim provides both values inside its driver-defined opaque config:

config:
  - requests: ["rdma-nic"]
    opaque:
      driver: networking.example.com
      parameters:
        apiVersion: networking.example.com/v1
        kind: NICConfig
        subnetID: subnet-a
        ibPKey: "0x8001"
        vlan: "100"           # not extracted (no CEL declared for it)

Alpha matching behavior:

  • If the device is clean, the first compatible claim sets the lock to:
    • subnet = subnet-a
    • pkey = 0x8001
  • A later claim whose extraction yields the same subnet and pkey may share the device.
  • A request whose extraction yields subnet = subnet-a but pkey = 0x8002 makes the device not a viable option because all declared keys must match.
  • A claim whose opaque config has no ibPKey field will cause object.ibPKey to error during CEL evaluation; this is a ResourceSlice authoring error and aborts scheduling for the Pod.
  • The vlan field is ignored because the driver did not declare a CEL expression for it. It flows through opaquely to the driver.

Test Plan

[x] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.

Prerequisite testing updates

Existing DRA scheduling tests should pass before adding sharing affinity tests.

Unit tests
  • pkg/scheduler/framework/plugins/dynamicresources: Coverage for affinity matching logic, including:
    • Filter: device with matching lock passes
    • Filter: device with conflicting lock is excluded
    • Filter: unlocked device with sufficient capacity passes
    • Filter: extraction yields no non-empty keys for the claim → device filtered out
    • Filter: claim with extra fields in opaque config beyond what the slice’s CEL extracts → extra fields ignored, device passes if extracted keys match
    • Filter: claim has multiple opaque configs producing conflicting same-key values across one or more extractors → device filtered out
    • Filter: CEL evaluation error (missing field, non-string return, cost budget exhausted) → scheduling aborts with a fatal error, no silent empty lock, no per-device Event spam
    • Filter: device with AffinityStates[deviceID].Status == AffinityStatusUnreconstructable is excluded for new sharing-affinity scheduling
    • First-fit verification: with two feasible devices on a node (one locked-compatible, one clean), allocator picks the first in ResourceSlice order (no affinity-aware preference; preference lands in beta)
    • Reserve: first claim sets lock; second claim with same extracted values succeeds
    • Reserve: second claim with conflicting extracted values fails
    • Unreserve: tentative lock is rolled back
    • Legacy claims with non-reconstructable affinity (no keys produced or CEL fails) cause the device to be marked unknown rather than establishing or joining a lock
    • Legacy-claim handling: all scenarios from the Handling Legacy Claims with Unreconstructable Affinity table
    • Compatibility matrix: device in a slice without sharingAffinity is unaffected — claims with or without matching opaque configs both pass (Legacy/Basic and Permissive Sharing rows)
    • Strict Gating: device’s pool has sharingAffinity but at least one extractor is not satisfied by the claim’s opaque configs (at least one extractor returns empty) → device filtered out
    • Multi-request scoping: claim with two requests (mgmt-nic and data-nic) each with distinct opaque configs → each request resolves independently; one request’s extracted values do not influence the other’s filter decision or lock state on a different device
  • staging/src/k8s.io/api/resource/v1: Coverage for the new ResourceSliceSpec.SharingAffinity field, including:
    • Validation: sharingAffinity exceeding max 8 entries is rejected
    • Validation: duplicate name values within a slice’s sharingAffinity list are rejected (declarative +listMapKey=name uniqueness)
    • Validation: CEL expressions are syntactically valid at admission (parse-check only; runtime cost is bounded per-eval at the scheduler)
    • Validation: a single ResourceSlice with both devices and sharingAffinity set is rejected (mutual exclusion)
    • Round-trip serialization of SharingAffinityExtractor
Integration tests
  • Affinity matching with multiple claims to same device: in a single-device topology, verify that a second compatible claim shares the locked device and extends the affinity lock. Tests Filter correctness and Reserve-phase lock extension; preference for the locked device when alternatives exist is a beta concern (see Affinity-aware scoring (planned for Beta) ).
  • Affinity mismatch causing allocation to different device
  • Affinity lock clearing when all claims release a device
  • Interaction with consumable capacity constraints (KEP-5075)
  • Scheduler restart: AffinityStates correctly reconstructed by re-running CEL extraction against the opaque configs of existing bound ResourceClaims; devices with non-reconstructable active claims (no keys produced, CEL eval failure) have AffinityStates[deviceID].Status == AffinityStatusUnreconstructable
  • Parallel scheduling: two Pods whose extracted affinity values conflict target the same device — one wins Reserve, the other is requeued
  • DRASharingAffinity disabled in the scheduler: devices in a pool that declares sharingAffinity are skipped, not treated as unconditionally shareable; pools without the field are unaffected
  • DRASharingAffinity toggled: enabling after claims exist does not disrupt already-bound workloads, and legacy in-use devices are conservatively filtered until clean
  • Invalid opaque config at scheduling time: regression test that malformed configs (no apiVersion/kind, unparseable parameters) cause CEL extraction to fail deterministically and the device is filtered out rather than crashing the scheduler
  • Permissive Sharing (slice has no AE): Slice without sharingAffinity, claim with arbitrary opaque config — verify scheduler allows the allocation and opaque config is not evaluated for affinity
  • Ghost Lock: Pod is Assumed (tentative lock set) but Bind fails — verify the lock is cleared immediately and the next Pod in the queue can claim the device with a different affinity value
  • Legacy Device Migration: 5 Pods are already running on NICs in a slice without sharingAffinity; the driver republishes the slice with sharingAffinity declared; a 6th Pod arrives with a matching opaque config — verify each affected device has AffinityStates[deviceID].Status == AffinityStatusUnreconstructable and the 6th Pod is filtered from those devices until all legacy claims drain
  • Partial Key: Pool declares two extractors, subnet and pkey. Claim’s opaque config has subnetID (so subnet resolves) but no ibPKey (so the pkey extractor returns empty and is not satisfied) — verify the device is filtered out
  • Multiple Extractors, All Required: Pool declares two extractors, vendor and subnet. A claim providing both vendor and subnetID — verify the device is eligible and the merged effective affinity map is {vendor, subnet}. The same claim missing vendor leaves the first extractor unsatisfied — verify the device is filtered out, confirming that satisfaction is required across every extractor in the pool, not just one.
  • Duplicate Key Across Metadata Slices (invalid pool): A pool publishes two metadata slices, each declaring an extractor with key subnet. Neither slice is rejectable at admission in isolation — verify the allocator’s pool validation marks the pool invalid and fails closed for every device in it, rather than silently picking one extractor’s value. After the driver republishes with the colliding key renamed, the next scheduling cycle succeeds.
  • Expression limits: an expression one byte over CELSelectorExpressionMaxLength (10 Ki) is rejected at admission and one at exactly the limit is accepted; a set of extractors whose combined estimated cost exceeds SharingAffinityCELMaxCost is rejected, both within one slice and across a pool’s metadata slices
  • First-Fit Behavior: Two devices available on a node, one already locked to subnet-X, one clean; new claim whose extraction yields subnet-X — verify the claim is allocated successfully (Filter excludes nothing; either device is feasible) and that the chosen device matches the allocator’s existing first-fit ordering. Alpha does not require the locked device to be preferred; affinity-aware preference is delivered in beta
  • Driver Backstop: Slice has no sharingAffinity, two claims with incompatible opaque config land on the same device — verify scheduler allows both (permissive), and NodePrepareResources rejects the incompatible claim
  • NodePrepareResources failure does not clear lock: Claim is bound and lock is set in the scheduler cache, but NodePrepareResources fails on the node — verify the affinity lock remains in the scheduler cache
  • Multi-request, multi-device, no cross-talk: Single claim with two requests (mgmt-nic, data-nic), each with a distinct opaque config yielding subnet=A and subnet=B targeting two different devices in slices that declare extraction — verify both requests succeed in the same scheduling cycle and each device locks to its own request’s extracted values without influence from the sibling request
  • Restart with inconsistent reconstructable locks: two active reconstructable claims on the same device produce conflicting extracted key values during restart reconstruction — verify the device is marked AffinityStates[deviceID].Status = AffinityStatusUnreconstructable, a warning is logged, and new sharing-affinity scheduling is blocked on that device until all claims drain
  • sharingAffinity mutation with active claims: slice republishes with renamed CEL keys or a new extractor; pre-existing claims now extract a different key set than the current LockedAffinity — verify the device has AffinityStates[deviceID].Status == AffinityStatusUnreconstructable after the next reconciliation/restart and is filtered out for new sharing-affinity scheduling until existing claims drain
  • CEL cost budget: pathological CEL expression intentionally exhausts the per-eval cost limit — verify the scheduler treats this as an extraction failure (device filtered out + Event), does not stall, and remains responsive to other scheduling work
e2e tests
  • End-to-end test with mock DRA driver publishing sharingAffinity
  • Multi-pod scheduling: Pods with matching extracted affinity values share the same device
  • Multi-pod scheduling: Pods with conflicting extracted affinity values are placed on different devices
  • Lock lifecycle: last Pod deleted → lock cleared → new Pod with different affinity value can claim the device
  • Rollout scenario: existing Pods running on devices in a slice with no sharingAffinity; driver republishes the slice with sharingAffinity; verify existing Pods continue running and new Pods respect the new constraint after legacy claims drain

Graduation Criteria

Alpha

  • Feature implemented behind a single feature gate DRASharingAffinity covering both the ResourceSlice.spec.sharingAffinity API field and the scheduler logic that runs CEL extraction and matches results against the lock
  • API field added to ResourceSlice (SharingAffinityExtractor on ResourceSliceSpec)
  • Scheduler runs CEL extraction over claim opaque configs to derive affinity keys; CEL evaluates in the standard ValidatingAdmissionPolicy environment with a per-evaluation cost budget
  • Scheduler Filter plugin enforces affinity matching
  • Scheduler tracks affinity in AllocatedState.AffinityStates
  • Unit and integration tests
  • Documentation for driver authors covering CEL expression authoring, multi-version guard idioms (e.g., object.apiVersion == "x/v1" && object.kind == "Foo" ? ... : ""), and rollout guidance
  • Alpha documentation explicitly calls out the lack of lock-breaking preemption semantics for incompatible locks
  • Alpha documentation explicitly calls out string-only affinity matching (alpha CEL expressions must return string)
  • Distinct alpha diagnostics emitted via scheduler logs and (best-effort) events for: (a) compatibility mismatch with the current lock, (b) no keys produced by any extractor for the claim, (c) CEL evaluation failure (including missing fields, non-string return, or cost-budget exhaustion), and (d) unknown lock state due to legacy / non-reconstructable / inconsistent active claims — sufficient to attribute a filtered scheduling decision without relying on metric labels alone

Beta

  • Gather feedback from DRA driver developers
  • Address any issues found in alpha
  • Affinity-aware scoring contribution: Extend DynamicResources.computeScore with a sharing-affinity term — prefer nodes where the target device is already locked to a compatible value (consolidation), then clean devices, with unknown-affinity devices deprioritized. Extend the structured-parameters allocator’s per-node device selection with the same preference. Follows the per-feature additive scoring pattern that Prioritized List (KEP-4816, shipped in 1.35) and Extended Resources (KEP-5004) already use; not blocked on the general-purpose scoring discussion in #4970 . See Affinity-aware scoring (planned for Beta) .
  • Observability of lock state: Surface effective per-device lock state for operators — exact mechanism TBD in beta. Candidates include scheduler-side metrics/gauges keyed by device and parameter hash, scheduler Events on Pending pods naming the locking claim, aggregation tooling over existing ResourceClaim.status.allocation, or a dedicated scheduler-owned API resource.
  • E2e tests stable
  • Performance validation with high pod churn

GA

  • At least 2 production drivers using sharing affinity
  • No significant issues reported
  • Conformance tests if applicable

Upgrade / Downgrade Strategy

Upgrade: Existing ResourceSlices without sharingAffinity continue to work. New field is additive. See the Compatibility Matrix for how the scheduler and driver behave across all combinations of slice sharingAffinity and claim opaque config presence.

This is a single-stage rollout: no claim-side API change is required, and workloads keep their existing opaque configs throughout.

The recommended sequence:

  1. Enable the gate (DRASharingAffinity) on apiserver and scheduler. Until any slice declares sharingAffinity, this is a no-op for scheduling.

  2. Publish sharingAffinity on ResourceSlices: add the CEL extractor block to the relevant slices. Strongly recommended to do this on idle slices first (see “Adding sharingAffinity to an in-use slice” below). From this point on, the scheduler locks a device when a claim targeting it produces a non-empty CEL extraction; for requests whose opaque configs produce no non-empty extractions, the device is not a viable option.

Why this order: when the DRASharingAffinity gate is OFF on the apiserver, sharingAffinity is stripped from incoming ResourceSlice writes (standard alpha-field handling). A driver that publishes the field before the gate is enabled will see it dropped at write time — reported by the ResourceSlice controller helper as a DroppedFieldsError, not silently. Enabling the gate later does not retroactively restore it, and the driver must republish.

Additional rollout consideration — workload-schema readiness: independent of step ordering, enabling extraction on slices before workloads have rolled to a compatible opaque-config schema may make the affected devices no longer viable options for those workloads on their next scheduling attempt, either landing them on non-extraction slices (capacity strand) or driving them to Pending if no non-extraction slices exist. The slice update should be coordinated with any required workload-side opaque-config evolution.

For the common case where the driver’s opaque config schema is unchanged and the slice’s CEL expressions simply pull existing fields out, no workload-side change is needed at all.

Minimizing capacity stranding: when adding sharingAffinity to slices for the first time, the ideal sequence is:

  1. Wait for affected devices to be idle (clean).
  2. Update the ResourceSlice to include the sharingAffinity block.
  3. Allow the scheduler to establish the first known lock with a new claim.

During mixed rollouts (some slices with sharingAffinity, some without), Strict Gating automatically routes claims whose opaque configs lack the declared keys onto non-extraction slices — extraction-slice devices are not viable options for those requests by design. The remaining gap is for claims whose opaque configs do produce the keys: in alpha they pass Filter on both extraction and non-extraction slices and the allocator has no preference between them, so capacity can strand on either side of the rollout. Affinity-aware preference is delivered in beta. For predictable rollout behavior, drain devices before adding sharingAffinity, as recommended above.

Adding sharingAffinity to an in-use slice: A driver may add or update sharingAffinity on a slice whose devices already have bound ResourceClaims. The scheduler handles this conservatively:

  • Pre-existing claims continue to run and are not evicted.
  • For each in-use device on the updated slice, the scheduler re-runs CEL extraction against every bound claim’s opaque config and populates LockedAffinity if all claims yield a consistent key map. If extraction fails or claims yield conflicting maps, the device is marked AffinityStatusUnreconstructable and excluded from new sharing-affinity placements until all its claims drain.
  • Devices on the slice with no bound claims are unaffected: they behave normally on next allocation.

Driver Upgrades and Schema Evolution: The opaque-config schema and the slice’s CEL extractors both evolve with driver releases. The recommended playbook depends on the kind of change:

Upgrade classRecommended action
Driver patch with no schema change, same CELNone. Driver republishes an identical slice; transparent to the scheduler.
Driver release with additive schema fields, same extracted keysNone. Existing CEL guards on apiVersion/kind continue to match; new fields are unread by the extractor and flow through to the driver.
Driver release adds v2 schema alongside v1Extend each extractor’s CEL with a v1 OR v2 apiVersion guard reading the same field path (the canonical multi-version idiom). Both old and new workloads extract identically. After workloads have migrated off v1, a later driver release drops the v1 guard.
Driver release with breaking schema rename (e.g., subnetID → subnet) — affects opaque-config field path only; declared key set unchangedRepublish the slice with an extractor that reads both paths into the same key, e.g. has(object.subnet) ? object.subnet : object.subnetID. Both old and new workloads extract identically; the legacy branch can be dropped in a later driver release once workloads have migrated, or kept indefinitely (the cost is one ternary in the CEL string). Existing locks remain valid — no Unreconstructable transitions.
Driver release with renamed extracted key (e.g., subnet → vpc-subnet) — affects the declared key set itselfThis is a sharingAffinity mutation: re-extraction against bound claims yields a different key map than the stored LockedAffinity. Devices with active legacy claims transition to AffinityStates[deviceID].Status == AffinityStatusUnreconstructable until they drain. Drivers should prefer to drain affected slices first (per the rollout sequence above) and stage the change to off-hours.
DaemonSet rolling upgrade across nodesEach node’s slice flips when its driver pod restarts. Transient heterogeneity across slices is bounded by the rollout window. Strict Gating plus Unreconstructable quarantine make the window safe but capacity-stranding; rolling-update maxSurge / maxUnavailable should be tuned with this in mind.
Workload-side schema lag (workloads behind driver)Drivers should keep multi-version-guard CEL through at least one workload rollout cycle. Otherwise workloads hit “no keys produced → Strict Gating → Pending” until they upgrade.
Driver pod removed or unhealthyExisting DRA behavior applies — the slice eventually becomes stale and is garbage-collected by the kubelet plugin manager. Devices in a stale slice are unschedulable regardless of sharingAffinity. No additional handling is introduced by this KEP.

Handling requests that produce no keys: The scheduler treats a slice with sharingAffinity as a protected resource. If a request’s opaque configs yield no non-empty value from any declared extractor’s CEL, no device in that slice is a viable option for the request. CEL evaluation errors (missing field, non-string return, cost-budget exhaustion) are treated the same way. API validation prevents most malformed CEL expressions from reaching the scheduler in the first place (parse-check at admission).

Downgrade: If the feature gate is disabled:

  • The apiserver strips the sharingAffinity field from writes that CREATE new ResourceSlices, so a slice freshly created with the gate off has no sharingAffinity.
  • The apiserver PRESERVES the sharingAffinity field on writes that UPDATE existing ResourceSlices where the field was already set on the prior object (ratcheting drop-on-disable). This prevents silent data loss across gate flips. Users may explicitly clear the field by setting it to null.
  • The scheduler skips every device in a pool that still carries sharingAffinity, rather than returning those pools to unconditional sharing. Drivers clear the field to make them allocatable again.
  • The DRA driver becomes the sole authority for enforcing hardware compatibility at NodePrepareResources once the field is cleared.

Version Skew Strategy

  • kube-apiserver: Must be upgraded first to accept the new sharingAffinity field on ResourceSlice.
  • kube-scheduler:
    • A scheduler that understands this feature enforces extraction, tracks AffinityStates, and may conservatively set AffinityStates[deviceID].Status = AffinityStatusUnreconstructable when effective affinity cannot be reconstructed.
    • A scheduler that has the code but runs with the gate off skips every device in a pool that declares sharingAffinity rather than allocating it unenforced.
    • A scheduler predating the feature does not recognize the field at all and cannot fail closed, so placement may be overly permissive and the DRA driver remains the final safety backstop during NodePrepareResources. This is the case the apiserver-first rollout order exists to keep short.
  • kubelet: No changes required; kubelet does not interpret sharingAffinity.
  • DRA driver:
    • Drivers publish ResourceSlices with the sharingAffinity field.
    • Drivers must continue validating actual hardware compatibility at prepare time, especially during skew where an older scheduler may not enforce affinity constraints.

During version skew, the main outcomes are permissive scheduling by an older scheduler or conservative filtering by a newer scheduler when affinity state cannot be reconstructed. Both are operationally safe as long as the driver continues rejecting incompatible prepare-time configurations.

Production Readiness Review Questionnaire

Feature Enablement and Rollback

How can this feature be enabled / disabled in a live cluster?
  • Feature gate
    • Feature gate name: DRASharingAffinity — gates ResourceSlice.spec.sharingAffinity and the scheduler logic that runs CEL extraction and enforces matching
    • Components depending on the feature gate: kube-apiserver, kube-scheduler
Does enabling the feature change any default behavior?

Not on its own. The feature gate is a precondition; behavior only changes for slices that explicitly opt in via the new sharingAffinity field, and slices without it behave exactly as before. Once a driver publishes slices that opt in, however, DeviceRequests evaluated against those slices will be subject to CEL extraction and affinity-lock matching at scheduling time without any change to the claim itself — that is the intended effect of the driver’s opt-in, but it means a cluster can see scheduling outcomes shift purely from a slice-side change. With the gate enabled, the timing of that shift is controlled by when drivers begin advertising sharingAffinity on their slices; with the gate disabled, the apiserver strips the field on writes that create new slices. Slices that already have sharingAffinity set when the gate flips off keep their data (ratcheting drop-on-disable — see the rollback question below).

Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?

Yes. Disabling DRASharingAffinity causes:

  • API server to strip the sharingAffinity field on writes that CREATE new ResourceSlices (writes succeed, field is not stored).
  • API server to PRESERVE the sharingAffinity field on writes that UPDATE slices where the field was already set on the prior object (ratcheting drop-on-disable; prevents silent data loss across gate flips). Users may explicitly clear the field by setting it to null.
  • Scheduler to skip every device in a pool that still declares sharingAffinity, rather than allocating it with enforcement off (see Feature Gates ).

Existing allocations continue to work. New allocations onto affinity-declaring pools stop until drivers republish those pools without sharingAffinity, so a complete rollback is: disable the gate in the scheduler, then have drivers clear the field.

What happens if we reenable the feature if it was previously rolled back?

The scheduler resumes enforcing sharingAffinity for future placement decisions. Existing allocations are not evicted.

ResourceSlices that ALREADY HAD sharingAffinity set when the gate was disabled retain the field (ratcheting drop-on-disable preserved the data across the gate flip). The scheduler picks up enforcement on those slices immediately. ResourceSlices that were newly CREATED while the gate was off will not have the field; drivers may republish those specific slices with sharingAffinity to opt them in.

While the gate was disabled the scheduler skipped affinity-declaring pools entirely, so it introduced no inconsistent co-location there. Devices can still need reconstruction: slices CREATED while the gate was off had the field stripped, so their pools were allocated permissively, and claims predating the feature carry no extractable keys. On reenable, the scheduler reconstructs lock state per device by re-running CEL extraction over each device’s active claims. A device is marked AffinityStatusUnreconstructable if the extracted key maps across its claims are inconsistent, or if extraction produces no non-empty keys at all. Such devices are excluded from new sharing-affinity placements until all active claims drain, after which they become clean and can establish a lock normally.

Are there any tests for feature enablement/disablement?

Yes, unit tests will cover the feature gate behavior for API validation and scheduler logic.

Rollout, Upgrade and Rollback Planning

How can a rollout or rollback fail? Can it impact already running workloads?

Rollout failure modes include:

  • Older scheduler after API enablement: a scheduler without the gate skips every device in a pool that declares sharingAffinity, so those devices are unschedulable until the scheduler is upgraded or drivers stop publishing the field.
  • Newer scheduler enabling conservative handling on legacy in-use devices: devices with non-reconstructable active claims may be filtered until they are clean, which can temporarily reduce effective schedulable capacity.

Rollback failure mode: if the scheduler is rolled back while the API server still serves the field, devices in affinity-declaring pools stop accepting new allocations until drivers clear sharingAffinity.

Running workloads are not evicted by this feature; the impact is on future placement decisions, not on already-running pods.

What specific metrics should inform a rollback?

Each signal below names a metric trajectory and the condition under which rollback beats a forward fix. Two levers exist and they are not interchangeable. Having drivers republish the affected pools without sharingAffinity is the per-pool lever: it restores permissive sharing immediately and needs no scheduler restart. Disabling the gate in the scheduler is the cluster-wide lever, but because a scheduler with the gate off skips affinity-declaring pools entirely, it must be paired with drivers clearing the field or those devices become unschedulable. Unless stated otherwise, “rollback” below means the driver-side lever.

  • sharing_affinity_unreconstructable_devices does not decay over an extended window (remains near its post-enablement peak after the expected workload-churn window). Indicates legacy in-use devices are not draining naturally and conservative handling will continue to suppress schedulable capacity indefinitely. Rollback restores permissive sharing for these devices; forward-fix has no lever.
  • sharing_affinity_lock_conflict_total rises for claims that operators expect to be compatible (cross-check against sharing_affinity_compatible_reuse_total staying flat). Indicates the driver’s CEL extractor is producing incorrect keys and over-filtering legitimate placements. Rollback re-enables fungible sharing while the extractor is corrected and re-published.
  • sharing_affinity_no_keys_extracted_total rises for workloads that previously placed successfully. Indicates extraction is silently dropping required parameters — typically claim opaque-config schema drift or extractor mis-targeting. Rollback restores placement while the driver re-publishes correct extractors.
  • Sustained increase in unschedulable DRA-backed pods beyond pre-enablement baseline (post-enablement P95 stays elevated over the cluster’s typical workload-churn window, with no corresponding workload increase). Indicates the combined effect of conservative handling and extractor-driven filtering is starving production placements faster than the feature’s benefits accrue.
  • Scheduler Filter/PreFilter latency regression measurable in scheduler_framework_extension_point_duration_seconds traceable to the affinity code path on DRA-heavy clusters. Rollback removes the per-claim CEL evaluation cost while the regression is profiled and fixed.
  • Scheduler panics or crashes in the affinity code path that cannot be patched in a short window. Rollback (disable the feature gate) is the lower-risk remediation than rolling an in-flight fix to production schedulers.

Note: rising rates of driver prepare-time rejections are not a rollback signal for this feature — those are exactly the failure mode KEP-5981 reduces. If they rise after enablement, the cause is upstream of the scheduler (driver bug, extractor bug) and rolling back would make them more frequent, not fewer.

Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?

Will be tested before beta.

Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?

No.

Monitoring Requirements

How can an operator determine if the feature is in use by workloads?

Two signals, in order of cost-to-collect:

  1. Metrics (recommended for fleet-scale visibility).

    • sharing_affinity_locked_devices reports how many devices the scheduler currently has under affinity lock. Any non-zero value confirms the feature is in active use in this cluster — the primary fleet-adoption signal.
    • A rising sharing_affinity_extractor_evaluations_total counter indicates the scheduler is actively running CEL extraction against published sharingAffinity slices, even if no locks have stuck yet.

    Both metrics are defined in the “Are there any missing metrics…” section below.

  2. Object inspection (single-cluster diagnosis). ResourceSlice objects that declare sharingAffinity show driver-side opt-in; ResourceClaims whose opaque configs yield non-empty keys under one of the declared extractors show workload-side participation. Useful for diagnosing a specific cluster, but not for fleet observability.

How can someone using this feature know that it is working for their instance?

There are no events or logs to show that the scheduler is excluding devices due to affinity locks, and adding them may be noisy. Instead, users can observe that the affinity is respected. A user should be able to observe that:

  • compatible claims are eligible to reuse already-locked devices (alpha does not actively prefer them; affinity-aware preference lands in beta),
  • when the scheduler has reconstructable affinity state, devices with incompatible affinity locks are not viable options for incompatible requests — gating happens before bind/prepare,
  • devices with unknown legacy affinity state are conservatively excluded until they become clean.

In practice, this should be visible through scheduler logs, scheduler events, and (where implemented) scheduler metrics.

What are the reasonable SLOs (Service Level Objectives) for the enhancement?

This enhancement should not materially regress baseline DRA scheduling latency for clusters that do not use sharingAffinity.

For clusters that do use the feature, the primary objective is correctness of compatibility-aware placement with bounded incremental scheduling overhead.

What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?

Useful SLIs include:

  • rate of scheduling attempts filtered due to sharing-affinity mismatch (no keys produced, CEL eval error, or lock mismatch),
  • rate of devices with AffinityStates[deviceID].Status == AffinityStatusUnreconstructable,
  • share of successful placements that reuse already-locked compatible devices,
  • prepare-time rejections by the DRA driver caused by incompatible hardware configuration.
Are there any missing metrics that would be useful to have to improve observability of this feature?

This feature would benefit from scheduler-observable counters, gauges, and/or events for:

  • sharing_affinity_locked_devices — number of devices currently holding at least one active affinity lock. The primary “feature in active use” signal: non-zero means the scheduler is actively maintaining affinity constraints in this cluster.
  • sharing_affinity_unreconstructable_devices — number of devices currently in conservative quarantine because their lock state cannot be reconstructed (legacy active claims, or extractors that no longer produce keys for those claims). If paired with sharing_affinity_locked_devices, they describe the cluster’s current device-state distribution under the feature.
  • sharing_affinity_extractor_evaluations_total — total CEL extractor evaluations the scheduler has performed. Useful as a “feature is being exercised” signal independent of placement outcomes. If this is rising but sharing_affinity_locked_devices stays at zero, the scheduler is evaluating extractors but no new locks are being established. Common causes, distinguishable by inspecting the other metrics in this list:
    • candidate devices are unreconstructable (sharing_affinity_unreconstructable_devices > 0) — legacy claims have not drained yet,
    • extractors do not match the claims’ opaque-config schemas (sharing_affinity_no_keys_extracted_total rising) — driver configuration or claim-template issue,
    • lock conflicts are blocking placements (sharing_affinity_lock_conflict_total rising) — competing claims with incompatible affinity values.
  • sharing_affinity_lock_conflict_total — candidate devices filtered out during Filter because the device’s existing affinity lock is incompatible with the claim’s extracted keys. This is the “feature is gating incompatible placements as designed” signal — non-zero is expected and healthy; the rate scales with claim diversity, not feature health.
  • sharing_affinity_no_keys_extracted_total — candidate devices filtered out during Filter because the slice’s extractors produced an empty effective affinity-key map for the claim. The adoption-gap signal: a high ratio of no_keys_extracted_total / extractor_evaluations_total means the feature is wired up but workloads aren’t engaging. Settles low in mature clusters; drops over time as workload teams update ClaimTemplates during rollout.
  • sharing_affinity_compatible_reuse_total — successful placements onto an already-locked device whose affinity values matched. It can stay zero when the feature is in use but the scheduler picks a different device for scoring reasons (spread, topology preference). Use sharing_affinity_locked_devices as the in-use indicator and sharing_affinity_compatible_reuse_total as the “preference is delivering value” indicator.

Counters above give cluster-wide visibility; individual pods that fail to schedule also need a clear per-pod reason. The scheduler should surface a specific diagnostic on the pod’s PodScheduled=False condition (and corresponding scheduling event) whenever an affinity check rejects a device, so the workload owner can tell why their pod is stuck. For example:

  • claim’s opaque configs produced no non-empty key under any sharingAffinity extractor on slice <slice> (device <id> filtered),
  • CEL expression <key> failed to evaluate against claim’s opaque config: <error> (device <id> filtered),
  • device <id> is locked to incompatible affinity values,
  • device <id> has unknown affinity state due to legacy or invalid active claims.

Dependencies

Does this feature depend on any specific services running in the cluster?
  • KEP-5075 (Consumable Capacity) for multi-allocatable devices

Scalability

Will enabling / using this feature result in any new API calls?

No new API calls. Affinity data is extracted from ResourceSlice and ResourceClaim objects already fetched by existing informers.

Will enabling / using this feature result in introducing new API types?

No. Only new fields on existing types.

Will enabling / using this feature result in any new calls to the cloud provider?

No.

Will enabling / using this feature result in increasing size or count of the existing API objects?
  • ResourceSlice: One or more additional metadata slices per pool that uses this feature (each carrying sharingAffinity and zero devices); the metadata slices in a pool jointly hold up to 8 extractors, each with up to 8 CEL expressions. Device-bearing slices are unchanged.
  • ResourceClaim: No change. The claim side carries no new fields.
Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?

Negligible. The Filter phase evaluates every extractor’s CEL against every opaque-config object on the claim. Worst-case CEL evaluation count per candidate device is O(extractors × opaque-configs), bounded by ≤8 extractors × N opaque configs (typically ≤2), and gated by the shared per-pool cost budget (SharingAffinityCELMaxCost, equal to the cost allowed for a single CEL selector). CEL programs are compiled at slice-write admission time and cached; per-evaluation cost is small and bounded. The lock-comparison itself is O(k) where k ≤ 8 — a map lookup per key.

Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?

No. The per-component impact is bounded:

  • Scheduler RAM: AffinityStates adds one map[string]string (up to 8 entries) per device with active affinity locks — proportional to active shared allocations, not total devices. A cluster with 1,000 actively-shared devices carries well under 1 MB of state; unshared devices contribute zero. Compiled CEL programs are cached cluster-wide keyed by expression string; the cache is bounded by ≤8 extractors per slice, with cross-slice deduplication in practice.
  • Scheduler CPU: Running CEL extraction during Filter is small per-candidate and bounded by the cost budget. Programs are compiled once and reused.
  • etcd disk: Slightly larger ResourceSlice objects (bounded by the caps above). ResourceClaim size is unchanged.
Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?

No.

Troubleshooting

How does this feature react if the API server and/or etcd is unavailable?

Like existing scheduler-driven DRA logic, this feature depends on informer state and cached API data. Temporary API server or etcd unavailability does not by itself invalidate already-computed in-memory lock state, but new pods will not be scheduled during unavailability. Sustained control-plane unavailability may delay reconciliation of claim release, slice updates, or restart reconstruction.

The driver remains the final enforcement authority at prepare time.

What are other known failure modes?

Known failure modes include:

  • No keys produced: the request’s opaque configs do not yield any non-empty value from at least one sharingAffinity extractor on the slice; the device is not a viable option for the request (Strict Gating outcome).
  • CEL evaluation failure: a CEL expression returns a non-string, references a missing field on the parsed opaque parameters, or exhausts the per-evaluation cost budget. ResourceSlice authoring error: the scheduler aborts allocation and fails scheduling for the Pod. The scheduler never silently establishes an empty lock.
  • Unreconstructable affinity state: the device has active allocations whose affinity cannot be reconstructed (legacy claims, extractors changed such that they no longer produce keys, or CEL failure during reconstruction), so it is conservatively filtered until clean.
  • Prepare-time driver rejection: despite scheduler filtering, the driver may still reject an incompatible or stale placement and that rejection is the final safety backstop.
  • Partial feature gate enablement: if the feature gate is enabled on the API server but not the scheduler (or vice versa), the sharingAffinity field may be persisted but not enforced, or enforced from cached objects but unable to be persisted on new writes. Ensure the gate is enabled on both kube-apiserver and kube-scheduler.
What steps should be taken if SLOs are not being met to determine the problem?

Recommended debugging flow:

  1. Inspect the relevant ResourceSlice and confirm it declares the expected sharingAffinity extractors (well-formed CEL expressions, expected key names).
  2. Inspect the ResourceClaim and confirm at least one of its opaque-config objects has the shape the slice’s CEL expressions expect (matching apiVersion/kind guards, populated fields).
  3. Mentally (or via kubectl + a CEL test harness) run each CEL expression against the claim’s parsed opaque parameters and verify it returns a non-empty string value.
  4. Check whether the target device is already locked to incompatible values (lock state is in the scheduler’s in-memory cache — check scheduler logs for filter reasons mentioning affinity mismatch).
  5. Check whether the device is being treated as having unknown affinity state because of legacy or non-reconstructable active claims.
  6. Review scheduler logs/events for explicit filter reasons (no keys produced, CEL eval failure, cost-budget exhaustion, lock mismatch, unknown state).
  7. If the scheduler allowed placement but the driver rejected prepare, inspect driver logs to determine whether the issue was stale scheduler state, unsupported config, or an actual device-level incompatibility.

Implementation History

  • 2026-03-27: Initial KEP issue created

  • 2026-03-30: KEP document drafted

  • 2026-04-28: Pivoted from a well-known JSON schema inside OpaqueDeviceConfiguration to a typed Structured sibling field on DeviceConfiguration, based on wg-device-management feedback that the scheduler should not interpret opaque payloads.

  • 2026-05-14: Pivoted again from the typed Structured sibling on DeviceConfiguration to driver-published CEL extraction on ResourceSlice.spec.sharingAffinity, per @johnbelamaric and @pohly review feedback on PR #5987. The claim side reverts to the status-quo opaque config (no new typed claim-side surface), and the scheduler’s only assumption becomes “find some GVK objects to which it can apply CEL.” Removed the DRAStructuredDeviceConfiguration feature gate (no longer needed). Consolidated into a single DRASharingAffinity gate.

  • 2026-05-16: Per @pohly review feedback (#discussion_r3246775761), dropped the GVK field from SharingAffinityExtractor. GVK matching is now the CEL author’s responsibility: a CEL expression guards on object.apiVersion and object.kind and returns the empty string when it does not apply to a given opaque-config object. The scheduler unions non-empty returns across all (extractor, opaque-config) pairs. Simplifies the API and lets one extractor handle multiple compatible config versions (e.g., v1beta1 + v1 with same field paths) without duplication.

  • 2026-05-17: Added a Driver Upgrades and Schema Evolution playbook under the Recommended Rollout section, covering schema-additive, multi-version-add, breaking-rename, key-rename, DaemonSet rolling upgrade, workload-side lag, and driver-pod removal cases. Codifies the multi-version-guard CEL idiom as the recommended upgrade primitive for breaking schema changes.

  • 2026-09-28: Addressed @pohly review feedback on PR #5987:

    • Replaced each entry’s cel map[string]string with a flat SharingAffinityExtractor{Name, Expression} list keyed by name (#discussion_r4120271856, #discussion_r3199293328).
    • Dropped the per-extractor Selector; pools using sharingAffinity are homogeneous, and heterogeneous pools became a rejected alternative (#discussion_r4120408066).
    • ResourceSlice authoring errors now abort allocation and fail scheduling for the Pod instead of skipping the device and emitting an Event (#discussion_r4120337098, #discussion_r4120605895).
    • A scheduler with the gate off now skips every device in a pool declaring sharingAffinity rather than ignoring the field (#discussion_r4120708639, #discussion_r4120870443).
    • Bounded each CEL expression by length and a shared per-pool cost budget — see Expression limits (#discussion_r4120767010).
    • AffinityStatus is now an int enum rather than a string; it is scheduler-internal cache state, and the zero value is AffinityStatusClean (#discussion_r4120810687).

Drawbacks

  • Adds a new cache dimension (AffinityStates) to the scheduler’s allocation tracking, increasing the surface area for reconstruction bugs on restart
  • Once a device is locked, its effective affinity cannot change until all claims on that device are released
  • Fragmentation risk remains if affinity values are too fine-grained
  • Conservative handling of legacy in-use devices can temporarily strand schedulable capacity during rollout or migration
  • If a pool declares sharingAffinity but a request’s opaque configs do not satisfy every extractor for a given candidate device (every extractor returning non-empty), that device is not a viable option for that request under the “Strict Gating” rule; if no device in the pool has all extractors satisfied by the request, the entire pool is filtered out. Drivers should coordinate with workload teams to ensure claims carry compatible opaque configs matching the published extractors before enabling sharingAffinity on the pool.
  • Per-evaluation CEL cost is bounded but non-zero; under pathological CEL expressions a slice could elevate per-candidate Filter cost. The per-eval cost budget and admission-time parse check are the primary mitigations.
  • Affinity locks are purely in-memory with no API or status field to inspect which devices are locked to which values. Debugging lock state in alpha requires scheduler logs; a future enhancement (tracked under Beta graduation) is to surface effective lock state via scheduler-side metrics, scheduler Events on Pending pods, aggregation over existing ResourceClaim.status.allocation, or a dedicated scheduler-owned API resource — exact mechanism TBD.

Alternatives

CEL selector on each extractor

SharingAffinityExtractor could carry an optional Selector *DeviceSelector — the same type used by DeviceRequest.selectors — restricting an extractor to the pool’s devices whose attributes satisfy the expression. Extractors without a selector would apply to every device. This enables heterogeneous pools: a networking driver publishing both NICs and VFs in one pool could bind subnet/pkey to NIC devices and parentNIC to VF devices, dispatching on existing attributes such as device.attributes["type"].

Sharing affinity instead applies to every device in a pool; a driver whose devices need different affinity keys publishes them in separate pools.

Rejected because:

  • No user story requires it: All three user stories (RDMA PKey alignment, FPGA bitstream sharing, single-subnet NIC sharing) are single-device-kind and need exactly one set of keys. Heterogeneous dispatch is a capability without a consumer.
  • Pool splitting is a complete substitute: A driver chooses its own pool names, so publishing NICs and VFs as two pools needs no API support and no scheduler involvement. The selector does not add heterogeneous affinity, only heterogeneous affinity within a single pool — a packaging convenience rather than a missing capability.
  • Uniqueness becomes a per-device property: Without selectors, name is unique across the whole list, so +listMapKey=name catches duplicates within a slice declaratively and the cross-slice check is a string-set comparison. With selectors, name can no longer be the list key — the same name may legitimately appear twice under disjoint selectors — so the declarative check is unavailable, and the authoritative check becomes “evaluate every selector against every device in the pool,” re-run whenever any slice in the pool changes. Both designs need allocator-side pool validation for the cross-slice case; the selector version makes that check O(extractors × devices) rather than O(extractors), and downgrades the diagnostic from “this pool is invalid” to “these devices in this pool are invalid.” See Key name uniqueness .

Per-device extractor reference

Instead of dispatching from the metadata slice, each Device could carry a field naming which sharing-affinity entries apply to it (for example sharingAffinityRefs []string matched against named extractors) — explicit dispatch, no CEL evaluation, readable directly from the device.

Rejected because:

  • It solves a problem the design does not have: Like the selector, it exists only to serve heterogeneous pools, which are out of scope by decision.
  • ResourceSlice size: The reference repeats on every device, in an object already capped at 128 devices and bounded by etcd limits, versus a single expression in the metadata slice.
  • A dangling-reference surface: A device could name an extractor that no metadata slice defines, which the API server cannot catch since it validates each slice in isolation — exactly the problem KEP-4815 documents for DeviceCounterConsumption references.

If that decision is ever revisited, the selector is the preferred mechanism: it reuses an existing type and adds no per-device bytes.

Well-known JSON schema inside OpaqueDeviceConfiguration

An earlier iteration of this KEP supplied claim-side affinity values via a scheduler-recognized JSON schema embedded in OpaqueDeviceConfiguration — i.e., a magic driver: resource.k8s.io opaque payload with apiVersion: resource.k8s.io/v1alpha1, kind: StructuredParameters that the scheduler would decode at Filter time.

Rejected because:

  • Violates the contract of Opaque: OpaqueDeviceConfiguration is, by name and design, opaque to the core API; only the driver is supposed to own its schema. A reserved global driver namespace decoded by the scheduler effectively turns part of the opaque payload into a core API surface without giving it API-server validation.
  • Weakly typed: A JSON-schema-inside-string approach is not strongly typed. API validation cannot enforce the schema, so most invariants shift to scheduler-side runtime decode/validation.
  • Constrained evolution path: Schema versioning lives inside an opaque blob rather than in the Kubernetes API surface.

Typed Structured sibling on DeviceConfiguration

A subsequent iteration of this KEP added a typed Structured sibling on DeviceConfiguration so the claim could carry affinity values in a strongly-typed, API-validated map:

type DeviceConfiguration struct {
    Opaque     *OpaqueDeviceConfiguration
    Structured *StructuredDeviceConfiguration   // proposed
}

type StructuredDeviceConfiguration struct {
    Requests   []string
    Parameters map[string]StructuredParameterValue
}

The scheduler would read Structured.Parameters directly to extract affinity values, with no need to interpret opaque payloads.

Rejected because:

  • New typed surface on a core API: Adds a new field to the resource.k8s.io API group that exists solely to serve scheduler-readable parameter extraction — a non-trivial API surface cost for a single consumer, when the underlying claim already carries the same data in driver-specific form via Opaque.
  • Forces claim duplication: For most realistic claims the affinity keys (subnet, pkey, vlan, …) are already present in the driver’s opaque config. The Structured field would require the claim author to duplicate those values in a second, scheduler-visible location, with no automatic enforcement that the two stay consistent.
  • Reviewer guidance (per @johnbelamaric, @pohly on PR #5987): The scheduler does not need a new typed claim-side field; it needs a way to extract affinity keys from the existing opaque config. The right mechanism for that extraction is driver-published CEL — which the driver already owns the schema for — placed on the ResourceSlice next to the device declarations.

Claim-side-only SharingAffinity (on DeviceRequest)

An alternative design adds a dedicated SharingAffinity field directly on DeviceRequest within ResourceClaim, with no corresponding declaration on the device or ResourceSlice. Sharing intent is expressed entirely from the consumer side — the device is mute on whether or how it can be shared:

type DeviceRequest struct {
    // ... existing fields ...
    SharingAffinity *SharingAffinity
}

type SharingAffinity struct {
    AffinityKey string
    Value       string
    Strategy    SharingStrategy
}

Rejected because:

  • Wrong layer: Sharing affinity is a property of how a device can be shared (a hardware-modal constraint declared by the driver), not a property of an individual request. Putting it on DeviceRequest would imply the consumer chooses the strategy, when in reality the device (and its driver) dictates which keys must agree across consumers.
  • Single key only: The shape above implies one AffinityKey/Value pair per request. Multi-key affinity (e.g., subnet + pkey + vlan) would require either repeating the field or introducing a list, both of which converge structurally to “a map of typed parameters” — which is exactly what the adopted CEL extraction produces, without requiring a new claim-side field at all.

Object Reference-based Affinity Matching

An alternative approach replaces inline affinity values with external object references. Instead of extracting values from the claim’s opaque config, the claim would reference a CRD (e.g., NetworkConfiguration) by name, and the device would declare which object kinds constrain sharing.

Rejected because:

  • Requires new fields on both ResourceClaim and Device (or ResourceSlice), whereas the adopted approach adds only a single slice-level sharingAffinity field.
  • Requires external CRD definitions, adding operational burden for cluster administrators.
  • Multi-dimensional affinity: A device may need affinity on multiple independent axes (e.g., subnet + VLAN). With object references, each axis would need its own CRD.
  • Indirect object references raise authorization and lifecycle concerns (who owns the CRD instance? what happens when it is deleted while claims reference it?).

CEL on runtime scheduler lock state (rejected variant)

A related CEL-based design — considered before the design pivot — would have had the driver publish a CEL expression on the ResourceSlice that evaluates whether a claim is compatible with the device’s current lock state:

sharingAffinity:
  lockExpression: >
    device.affinityLock['subnet'] == '' ||
    device.affinityLock['subnet'] == claim.AffinityValues['subnet']

Rejected because:

  • device.affinityLock is runtime scheduler state, not a static device attribute. Exposing it in CEL requires extending the evaluation context to include the scheduler’s in-memory AllocatedState, which breaks the current model where CEL evaluates against the ResourceSlice snapshot.
  • CEL expressions are powerful but opaque to the scheduler — it cannot extract which keys constrain sharing or what values to record in AllocatedState. The scheduler would need to both evaluate the expression AND separately track lock state, duplicating logic.
  • A circular variant — where one claim’s eligibility depends on another claim’s allocation — produces non-deterministic results depending on evaluation order.

Future Enhancements

The following ideas are out of scope for alpha but are worth exploring in beta/GA based on real-world feedback:

Affinity-aware scoring (planned for Beta)

This KEP’s Filter phase is sufficient for correctness — an incompatible locked device is filtered out, and any remaining candidate produces a valid allocation. It is not, however, sufficient for packing, at either scope:

  • Cross-node: stock Kubernetes scorers do not factor in sharingAffinity when scoring nodes yet. A subnet=X claim is just as likely to land on a node with a clean device as on a node that already has a compatibly-locked device.
  • Within-node: once a node is selected, the DRA structured-parameters allocator picks the first feasible device (first-fit). Among multiple feasible devices on the chosen node — say, one already locked to a compatible affinity value and one clean — there is no preference logic; the allocator may pick whichever appears first in the ResourceSlice.

The DynamicResources plugin already implements scoring for DRA (shipped in K8s 1.35). Prioritized List (KEP-4816) and Extended Resources (KEP-5004) already contribute their own additive terms to computeScore. This is the canonical extension point for new DRA features that need scoring.

KEP-5981 plans to add a sharing-affinity term to computeScore as a beta deliverable, following the same per-feature additive pattern. The broader “general-purpose DRA scoring” discussion continues in kubernetes/enhancements#4970 , but this KEP’s contribution does not block on a unified framework or on #4970 producing a generic API. Doing per-feature scoring here is consistent with how Prioritized List and Extended Resources already shipped scoring, and avoids stranding sharing affinity behind a multi-feature design effort.

The detailed score terms, weights, and tie-breakers will be designed and proposed as part of the beta graduation. Alpha intentionally does not commit to specific scoring shape so that beta has freedom to incorporate operational feedback from alpha.

Within-node device selection is addressed similarly by extending the allocator’s per-node selection logic so that, among feasible devices for a chosen sub-request, the allocator prefers a locked-compatible device over a clean one. This is a localized change to the structured-parameters allocator and is also a beta deliverable.

Alpha contract: correctness only. Packing of any kind (within-node or cross-node) is best-effort first-fit and is documented as a known limitation in Risks and Mitigations .

Priority-based Lock Preemption

Removed in PR review (2026-05): lock-breaking preemption is not a KEP-5981 deliverable. See Composition with DRA Preemption (KEP-5690) under Notes/Constraints/Caveats.

SharingStrategy (CanSetLock / NeverSetLock)

Alpha intentionally does not let claims control whether they may establish a new lock on a clean device. Any compatible claim can set the initial lock, and subsequent compatible claims can then reuse that locked device.

A future enhancement could add an explicit SharingStrategy on the claim side to control lock-setting behavior. Two candidate strategies are:

  • CanSetLock (default): The claim may land on a clean device and establish the lock. This matches the alpha behavior.
  • NeverSetLock: The claim may only be allocated to a device that already has a matching lock established by another claim. This is useful for background or batch jobs that should never consume a clean device and potentially fragment capacity. Caveat: NeverSetLock is a follower-only strategy — it requires at least one CanSetLock claim to establish the lock first. If no device is locked to the requested value, a NeverSetLock pod will remain unschedulable indefinitely. Implementations should document this dependency clearly and consider surfacing a scheduling event when a pod is blocked waiting for a lock that no leader has established.

If introduced in beta or later, the scheduler would evaluate this policy before capacity and key matching for unlocked devices. A claim with NeverSetLock would reject an unlocked device immediately, then continue searching for an already-locked compatible device.

This is deferred from alpha to keep the initial scope focused on the core problem: driver-declared sharing constraints plus scheduler-enforced lock tracking via CEL-extracted parameters.

Soft / Preferred Affinity Keys

The Alpha design enforces hard all-or-nothing matching across the pool’s extractors: every extracted key must agree with the device’s current lock, or the device is filtered out. Real-world hardware may have hierarchical constraints where some keys are strict sharing requirements (e.g., Subnet) and others are scheduling preferences (e.g., Traffic-Class or bandwidth profile).

A future enhancement could let the driver mark individual CEL expressions as required vs preferred:

  • required (default): Mismatch → device filtered out; key contributes to the lock (current behavior).
  • preferred: Mismatch → device passes Filter but is deprioritized in device selection; key does not contribute to the lock.

Infrastructure Needed

None