KEP-5491: DRA: List Types for Attributes
KEP-5491: DRA: List Types for Attributes
- Release Signoff Checklist
- Summary
- Motivation
- Proposal
- Design Details
- Production Readiness Review Questionnaire
- Implementation History
- Drawbacks
- Alternatives
- Infrastructure Needed (Optional)
Release Signoff Checklist
Items marked with (R) are required prior to targeting to a milestone / release.
- (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
- (R) KEP approvers have approved the KEP status as
implementable - (R) Design details are appropriately documented
- (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
- e2e Tests for all Beta API Operations (endpoints)
- (R) Ensure GA e2e tests meet requirements for Conformance Tests
- (R) Minimum Two Week Window for GA e2e tests to prove flake free
- (R) Graduation criteria is in place
- (R) all GA Endpoints must be hit by Conformance Tests within one minor version of promotion to GA
- (R) Production readiness review completed
- (R) Production readiness review approved
- “Implementation History” section is up-to-date for milestone
- User-facing documentation has been created in kubernetes/website, for publication to kubernetes.io
- Supporting documentation—e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes
Summary
The Device Resource Assignment (DRA) API currently allows scalar attribute values to describe device characteristics. However, many real-world device topologies require representing sets of relationships (e.g., multiple PCIe roots, NUMA nodes). This KEP introduces support for list-typed attributes in ResourceSlice and extends(redefine) ResourceClaim’s constraints[].{matchAttribute, distinctAttribute} semantics to fit both list-type attributes and primitive attributes supported previously.
Motivation
The ResourceSlice API allows users to attach scalar attributes to devices. These can be used to allocate devices that share common topology within the node. For certain types of topological relationships, scalar values are insufficient. For example, a CPU may have adjacency to multiple PCIe roots. This enhancement proposes allowing attributes to be lists. The semantics of the MatchAttribute and DistinctAttribute constraints must adapt to the possibility of lists. For example, rather than defining an attribute “match” as equality, it would be defined as a non-empty intersection, treating scalars as single-element lists. Conversely, “distinct” attributes for lists would be defined as an empty intersection.
Goals
- Support typed-list in device attribute values.
- Extends(redefine) the semantics of
ResourceClaim’sconstraints[].{matchAttribute,distinctAttribute}fields as below so that it can work with list-type attribute valuesmatchAttribute: it is defined as non-empty intersectiondistinctAttribute: it is defined as pairwise disjoint- note: scalar values are treated as single-element lists
- Keep monotonicity in constraint.
- Currently
Allocator’s algorithm assumes monotonic constraints only. Monotonic means that once a constraint returns false, adding more devices will never cause it to return true. This allows to bound the computational complexity for searching device combinations which satisfies the specified constraints. This KEP focuses to keep monotonicity ofmatchAttribute/distinctAttributesemantics.
- Currently
- Maintain backward compatibility and inter-operability for scalar-only attributes.
matchAttribute/distinctAttribute: existing constraint can work because scalar values are treated as single-value list- CEL expressions in device selectors: when the attribute type is updated, existing CEL won’t failed to compile. But, we will provide some type-agnostic helper function to achieve easier migration for users/DRA driver developers.
Non-Goals
- Introducing generic or complex boolean logic in constraints(KEP-5254: DRA: Constraints with CEL).
- Forcing all drivers to use list attributes immediately.
Proposal
The proposal has mainly two parts:
- Add list-types in
DeviceAttributeso that DRA drivers can expose the attribute values in typed list(int,string,boolean,version) - Extends the semantics of
MatchAttribute/DistinctAttributefield inDeviceConstraint- For
MatchAttribute:- Previously: it matches when the attribute values among candidate devices are identical (i.e.
∀i,j, v_i = v_j) - This KEP: it matches when the intersection (as a set) of all the list values among candidate devices is non-empty(i.e.
(∩ v_k != ∅))
- Previously: it matches when the attribute values among candidate devices are identical (i.e.
- For
DistinctAttribute- Previously: it matches when all the attribute values among candidate devices are distinct (i.e.
∀i,j, s.t. i != j, v_i != v_j) - This KEP: it matches when all the list values among candidate devices are pairwise disjoint (i.e.
∀i,j, s.t. i != j, v_j ∩ v_k = ∅)
- Previously: it matches when all the attribute values among candidate devices are distinct (i.e.
- For
API Changes
Introduce typed-list in DeviceAttribute
kind: ResourceSlice
spec:
devices:
- name: typed-list-attributes
attributes:
list-of-string:
strings: ["pci0000:00", "pci000:01"]
list-of-int:
ints: [0, 1, 2]
list-of-bool:
bools: [true, false, true]
list-of-version:
versions: ["1.0.0", "1.0.1"]
Introduce .includes function in CEL
When the attribute type was changed from scalar to list. Existing CEL won’t compile due to type mismatch.
// This CEL won't compile if attributes["foo"] type is changed from 1 (scalar) to [1](https://raw.githubusercontent.com/kubernetes/enhancements/master/keps/sig-scheduling/5491-dra-list-types-for-attributes/list)
attributes["foo"] == 1
To maintain backward compatibility for existing CEL expressions, it might be possible to override comparison operators (==, etc.) that allows for a list type where attributes["foo"] == 1 is equivalent to attributes["foo"] == [1]. But we don’t do this way because it wouldn’t be idiomatic and would diverge from normal CEL type system expectations and feels confusing to anyone that already has an understanding of how the CEL type system is suppose to work.
Instead, although user needs to rewrite the existing CEL expressions, we provide a helper function, .includes, which works in a type-agnostic way to make the CEL migration easier:
// assume attribute["foo"] is 1
attribute["foo"].includes(1) --> true
// assume attribute["foo"] is [1]
attribute["foo"].includes(1) --> true
.includes (<dyn>.includes(<dyn>)) has since been implemented not as a DRA-specific helper, but migrated into the shared, versioned CEL extension library k8s.io/apiserver/pkg/cel/library (library.Lists(library.ListsVersion(1)), available in the CEL environment from version 1.37 onward: kubernetes/kubernetes#140016). This means the function is reusable by any CEL-consuming API in Kubernetes, not just DRA device selectors.
During Alpha it was proposed to restrict .includes to the attributes[x] target only (kubernetes/kubernetes#137901). That was decided against and the issue closed: once the function is scoped under the versioned lists library it is a legitimate generic helper whose availability is governed by that library’s version, so a scalar receiver such as 1.includes(1) is acceptable and no DRA-specific restriction is applied.
Additionally, when re-evaluating a CEL expression that was already persisted (e.g. a selector saved before a downgrade, or evaluated by an n-1 component), list-typed attributes and .includes remain usable even if DRAListTypeAttributes is disabled at evaluation time: kubernetes/kubernetes#139395. This closes a version-skew gap identified during alpha (a previously-saved selector referencing list attributes could otherwise fail to re-evaluate on a gate-disabled or older component).
User Stories (Optional)
Story 1: Hardware Topological Aligned CPUs & GPUs & NICs
Assume several DRA drivers exposed device attribute resource.kubernetes.io/pcieRoot:
apiVersion: resource.k8s.io/v1
kind: ResourceSlice
metadata:
name: cpu
spec:
driver: "cpu.example.com"
pool:
name: "cpu"
resourceSliceCount: 1
nodeName: node-1
devices:
- name: "cpu-0"
attributes:
resource.kubernetes.io/pcieRoot:
strings:
- pci0000:01
- pci0000:02
- name: "cpu-1"
attributes:
resource.kubernetes.io/pcieRoot:
strings:
- pci0000:03
- pci0000:04
---
apiVersion: resource.k8s.io/v1
kind: ResourceSlice
metadata:
name: gpu
spec:
driver: "gpu.example.com"
pool:
name: "gpu"
resourceSliceCount: 1
nodeName: node-1
devices:
- name: "gpu-0"
attributes:
# Assume this driver is a bit old that keeps exposing string for the attribute
resource.kubernetes.io/pcieRoot:
string: pci0000:01
---
apiVersion: resource.k8s.io/v1
kind: ResourceSlice
metadata:
name: nic
spec:
driver: "nic.example.com"
pool:
name: "nic"
resourceSliceCount: 1
nodeName: node-1
devices:
- name: "nic-0"
attributes:
# Assume this driver is a bit old that keeps exposing string for the attribute
resource.kubernetes.io/pcieRoot:
string: pci0000:01
Then, user can create ResourceClaim resource which requests PCIe topology aligned CPU & GPU & NIC triple like below:
apiVersion: resource.k8s.io/v1
kind: ResourceClaim
spec:
requests:
- name: "gpu"
exactly:
deviceClassName: gpu.example.com
count: 1
- name: "nic"
exactly:
deviceClassName: nic.example.com
count: 1
- name: "cpu"
exactly:
deviceClassName: cpu.example.com
count: 2
constraints:
# "gpu-0", "nic-0" and "cpu-0" above can match
# because
# - "pci0000:01" is common.
# - string attribute can be treated as a single value list
- requests: ["gpu", "nic", "cpu"]
matchAttribute: k8s.io/pcieRoot
Story 2: Cross-driver NUMA Co-placement via a Standard Attribute
KEP-6072, which defines the standard device attribute resource.kubernetes.io/numaNode and is stable as of v1.37, is the first in-tree consumer of the list types proposed here. It shows the intended usage pattern end to end:
- A driver publishes
resource.kubernetes.io/numaNodeas anintslist: the device’s physical NUMA node (from the kernel’snuma_nodesysfs entry) first, followed by the same-socket NUMA nodes at the minimum ACPI SLIT distance. A driver that cannot compute the SLIT-based set may publish the physical node alone as a scalarint. - A workload asks for a GPU and a NIC served by two different drivers, and ties them together with a single
matchAttribute: resource.kubernetes.io/numaNode. - Because
matchAttributeis evaluated as non-empty set intersection and a scalar is treated as a single-element set, the two devices match whenever their NUMA sets overlap: a scalar4from one driver matches a list[4, 5, 6, 7]from another. Mixed scalar/list forms across drivers need no special handling from the user or from either driver.
Without list types, every driver would have to agree on one encoding for “the set of acceptable NUMA nodes” (for example a formatted string), and matchAttribute could only compare those encodings for exact equality – which fails precisely in the mixed scalar/list case that occurs while drivers are being rolled out.
KEP-6080 (derivedAttributes, targeting Beta in v1.38) is a second consumer: its CEL expressions read and return list-typed attributes and inherit the intersection semantics defined here.
Notes/Constraints/Caveats (Optional)
Risks and Mitigations
- Risk 1: Driver adoption lag
- Mitigation: scalar is treated as single value list
- Risk 2: Scheduler performance overhead
- bound lengths of the list-typed attribute values
Design Details
Go Type Definitions
DeviceAttribute
Note: The total number of individual attribute values per device (scalar fields plus all list elements combined) is limited to 48 (referring ResourceSliceMaxAttributeValuesPerDevice).
When any device in a ResourceSlice uses this feature or other advanced features such as taints, the ResourceSlice will be limited to at most 64 devices (referring ResourceSliceMaxDevicesWithAdvancedFeatures).
type DeviceAttribute struct {
...
// IntValues is a non-empty list of numbers.
//
// This is a beta field and requires enabling the DRAListTypeAttributes feature gate.
//
// +optional
// +listType=atomic
// +k8s:listType=atomic
// +k8s:beta(since: "1.38")=+k8s:optional
// +k8s:beta(since: "1.38")=+k8s:unionMember
// +featureGate=DRAListTypeAttributes
IntValues []int64 `json:"ints,omitempty" protobuf:"varint,6,opt,name=ints"`
// BoolValues is a non-empty list of true/false values.
//
// This is a beta field and requires enabling the DRAListTypeAttributes feature gate.
//
// +optional
// +listType=atomic
// +k8s:listType=atomic
// +k8s:beta(since: "1.38")=+k8s:optional
// +k8s:beta(since: "1.38")=+k8s:unionMember
// +featureGate=DRAListTypeAttributes
BoolValues []bool `json:"bools,omitempty" protobuf:"varint,7,opt,name=bools"`
// StringValues is a non-empty list of strings.
// Each string must not be longer than 64 characters.
//
// This is a beta field and requires enabling the DRAListTypeAttributes feature gate.
//
// +optional
// +listType=atomic
// +k8s:listType=atomic
// +k8s:beta(since: "1.38")=+k8s:optional
// +k8s:beta(since: "1.38")=+k8s:unionMember
// +k8s:beta(since: "1.38")=+k8s:eachVal=+k8s:maxBytes=64
// +featureGate=DRAListTypeAttributes
StringValues []string `json:"strings,omitempty" protobuf:"bytes,8,opt,name=strings"`
// VersionValues is a non-empty list of semantic versions according to semver.org spec 2.0.0.
// Each version string must not be longer than 64 characters.
//
// This is a beta field and requires enabling the DRAListTypeAttributes feature gate.
//
// +optional
// +listType=atomic
// +k8s:listType=atomic
// +k8s:beta(since: "1.38")=+k8s:optional
// +k8s:beta(since: "1.38")=+k8s:unionMember
// +featureGate=DRAListTypeAttributes
VersionValues []string `json:"versions,omitempty" protobuf:"bytes,9,opt,name=versions"`
}
Implementation (for evaluating constraints)
Since non-empty intersection constraint is monotonic, we would not need updating Allocator.Allocate() algorithm and can keep using constraint interface. We will just extend the current matchAttributeConstraint and distinctAttributeConstraint instances. Or, we could introduce constraint instances for proposed modes (e.g., nonEmptyIntersectionMatchAttributeConstraint, etc.).
Update (Beta): at Alpha, list-attribute constraint evaluation is implemented only in the structured/internal/experimental allocator variant; at Beta, it also rotates into internal/incubating, following the standard alpha/beta/GA allocator rotation model.
Test Plan
[x] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.
Prerequisite testing updates
None.
Unit tests
API validation:
- valid and invalid list-typed
DeviceAttributevalues (ints/bools/strings/versions) - exactly one value field is set per attribute, and every list is non-empty
- the per-device limit on the total number of attribute values, and the maximum length of string and version values
- feature gate transition behavior for the new fields on create and update
- valid and invalid list-typed
CEL:
- compilation and evaluation of device selectors over list-typed attributes
.includesapplied to both scalar and list-typed attributes- expressions referencing list-typed attributes stay evaluable in the stored expressions environment while the feature gate is disabled
Scheduler allocator:
matchAttributeallocates devices whose attribute values have a non-empty intersection, and rejects devices whose values are disjointdistinctAttributeallocates devices whose attribute values are pairwise disjoint, and rejects devices whose values overlap- scalar and list-typed values match against each other, a scalar being treated as a single-element set
- the intersection is restored when the allocator backtracks
k8s.io/dynamic-resource-allocation/cel:2026-09-13-93.2%k8s.io/apiserver/pkg/cel/environment:2026-09-13-76.8%k8s.io/dynamic-resource-allocation/structured/internal/experimental:2026-09-13-94.9%k8s.io/dynamic-resource-allocation/structured/internal/incubating:<TBD date>-<TBD coverage>k8s.io/kubernetes/pkg/apis/resource/validation:2026-09-13-97.0%k8s.io/kubernetes/pkg/registry/resource/resourceslice:2026-09-13-88.4%k8s.io/kubernetes/pkg/scheduler/framework/plugins/dynamicresources:2026-09-13-86.9%
Integration tests
test/integration/dra: the feature gate is exercised as part of the aggregate feature gate test group. Add allocation tests for list-typedmatchAttributeanddistinctAttributeconstraints, including device sets that mix scalar and list-typed values.
e2e tests
test/e2e/dra: add tests that a pod is scheduled or stays pending according to list-typedmatchAttribute(non-empty intersection) anddistinctAttribute(pairwise disjoint) constraints, including device sets that mix scalar and list-typed values.
Graduation Criteria
Alpha
- Feature implemented behind a feature flag (
DRAListTypeAttributes). The Feature gate is disabled by default. - Documentation provided
- Initial unit and integration tests
Beta
- Feature Gate
DRAListTypeAttributesis enabled by default. - List-typed
matchAttribute/distinctAttributeallocation is implemented in theincubatingallocator, which is the default implementation (see Implementation). - Integration and e2e tests for list-typed
matchAttribute/distinctAttributeallocation are in place (see Test Plan). - All the issues (https://github.com/kubernetes/kubernetes/issues/137905) which was identified in the initial implementation should be resolved.
- No major outstanding bugs.
- 1 example of real-world use case.
- Feedback collected from the community (developers and users) with adjustments provided, implemented and tested.
- All Beta PRR questions answered.
GA
- 2 examples of real-world use cases.
- Allowing time for feedback from developers and users.
Upgrade / Downgrade Strategy
- Upgrade: All fields introduced by this KEP are optional, so existing
ResourceSliceandResourceClaimobjects are unaffected. Drivers can start publishing list-typed attributes at any time; until one does,matchAttribute/distinctAttributebehave as before, because a scalar is treated as a single-element set. - Downgrade: Disabling
DRAListTypeAttributesprevents writing new list-typed attribute values. Already-stored values are preserved, but are not considered bymatchAttribute/distinctAttributeevaluation, so claims constraining on such an attribute fail to allocate.
Version Skew Strategy
For upgrade, existing ResourceClaim/ResourceSlice will still work as expected, as the new list-typed attribute fields are missing there.
For downgrade/skew: list-typed attribute values already stored in a ResourceSlice remain in etcd and are served as-is by kube-apiserver; they are not deleted or rewritten. However, if kube-scheduler is n-1 (or the gate is disabled on it) and a ResourceClaim’s constraints[].{matchAttribute,distinctAttribute} references an attribute that is list-typed, that scheduler does not read the list-typed values for constraint evaluation, so it cannot find a device satisfying the constraint. Allocation simply fails for that claim, leaving its pod Pending/unschedulable — no error, crash, or data loss, just an allocation that can’t succeed until the scheduler is upgraded (or the gate is re-enabled).
For version skew specifically involving CEL device selectors: .includes and list-typed attributes remain usable when re-evaluating an already-persisted CEL expression, regardless of the DRAListTypeAttributes gate state at evaluation time. So a selector expression referencing a list-typed attribute, once compiled while the gate was enabled, remains re-evaluable by an n-1 kube-apiserver or kube-scheduler that has the gate disabled, avoiding a hard failure on a previously-saved selector during a rolling downgrade.
Production Readiness Review Questionnaire
Feature Enablement and Rollback
How can this feature be enabled / disabled in a live cluster?
- Feature gate (also fill in values in
kep.yaml)- Feature gate name:
DRAListTypeAttributes - Components depending on the feature gate: kube-apiserver, kube-scheduler
- Feature gate name:
- Other
- Describe the mechanism:
- Will enabling / disabling the feature require downtime of the control plane?
- Will enabling / disabling the feature require downtime or reprovisioning of a node?
Does enabling the feature change any default behavior?
Basically, no. Just introducing new API fields in ResourceSlice which does NOT change the default behavior when any device attribute type was NOT changed.
However, please note that ResourceClaim’s matchAttribute/distinctAttribute semantics are CHANGED when some device attribute type are changed from scalar to list: matchAttribute requires a non-empty intersection and distinctAttribute requires pairwise disjointness. And, existing CEL device selectors comparing such an attribute with == will NOT compile any more. They need to be rewritten with .includes (see API Changes).
Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?
Yes. When disabled, DeviceAttribute with list-type values can no longer be created. Already-stored list-type attribute values are not deleted and are still served via the API as-is; they are simply not read by matchAttribute/distinctAttribute constraint evaluation while the gate is disabled. So if the attribute referenced by matchAttribute/distinctAttribute is list-typed, allocation for claims using that constraint will fail (see Version Skew Strategy).
This differs from CEL-based device selectors (e.g. DeviceClassSelector): a selector already referencing a list-typed attribute via .includes continues to evaluate correctly even after the gate is disabled, since list-typed attributes and .includes remain usable when re-evaluating an already-persisted CEL expression regardless of gate state. Only new selectors referencing a list-typed attribute cannot be created while the gate is disabled, symmetric to DeviceAttribute itself.
What happens if we reenable the feature if it was previously rolled back?
DeviceAttribute with list-type values can be created again, and the non-empty-intersection/pairwise-disjoint semantics for matchAttribute/distinctAttribute in ResourceClaim are available again.
Are there any tests for feature enablement/disablement?
Yes. The feature enablement and disablement coverage summarized in Unit tests is implemented in the following tests:
- The
ResourceSliceREST strategy tests exercise creation with the feature gate enabled and disabled and updates, including preservation of existing list-typed attributes after the gate is disabled. - The CEL tests verify that stored expressions with list-typed attributes remain valid when the feature gate is disabled.
- The allocator tests verify the disabled-gate behavior for list-typed
matchAttributeanddistinctAttributeconstraints.
Rollout, Upgrade and Rollback Planning
How can a rollout or rollback fail? Can it impact already running workloads?
This feature is implemented only in kube-apiserver and kube-scheduler. There is no node/kubelet component.
During a rollout of an HA control plane, some kube-apiserver instances may have the gate enabled while others do not. Then, a ResourceSlice write carrying list-typed attribute values is rejected when it lands on a gate-disabled instance, until the rollout completes cluster-wide. Values which are already stored are kept as-is.
Already running workloads are NOT impacted. Already allocated ResourceClaims are not affected by flipping the gate. Only re-allocation is affected, and only for claims whose constraints reference a list-typed attribute.
What specific metrics should inform a rollback?
Operators should watch for an increase in:
scheduler_unschedulable_pods{plugin="DynamicResources"}— an unexpected rise may indicatematchAttribute/distinctAttributeconstraints over list-typed attributes are failing to allocate as expected.scheduler_plugin_execution_duration_seconds{plugin="DynamicResources"}— a latency increase may indicate the additional list/set-intersection computation in the allocator’s constraint evaluation is more expensive than anticipated for the cluster’s attribute-list sizes.apiserver_request_total{resource="resourceslices"}/apiserver_request_total{resource="resourceclaims"}with non-2xx response codes — an increase may indicate validation/admission issues with list-typed attribute fields.
If any of these metrics show a sustained regression after enabling the feature, disabling the DRAListTypeAttributes feature gate is the expected rollback action.
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
This will be done manually before the Beta release by bringing up a KinD cluster and changing the feature gate for kube-apiserver and kube-scheduler individually, and the results will be documented here.
Handling of the new fields when the feature gate gets disabled is covered by unit tests in pkg/registry/resource/resourceslice and k8s.io/dynamic-resource-allocation/cel.
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?
No. This feature is purely additive: new optional fields on DeviceAttribute, and a semantics extension (not a field rename) for the existing matchAttribute/distinctAttribute constraint fields. No existing API, field, or flag is deprecated or removed.
Monitoring Requirements
How can an operator determine if the feature is in use by workloads?
Check for ResourceSlice objects whose devices set one of the list-typed DeviceAttribute fields (ints/bools/strings/versions), or ResourceClaim objects whose constraints[].{matchAttribute,distinctAttribute} reference such an attribute. There is no dedicated “feature in use” gauge metric; this is a reasonable API-inspection fallback since the fields themselves are the signal of use.
How can someone using this feature know that it is working for their instance?
- Events
- Event Reason:
FailedScheduling - Details: When no device combination satisfies the list-typed
matchAttribute/distinctAttributeconstraints, the scheduler records aFailedSchedulingevent on the Pod. The event message includes the allocation failure reported by theDynamicResourcesplugin, for example,cannot allocate all claims.
- Event Reason:
- API .status
- Condition name: N/A
- Other field:
ResourceClaim.Status.Allocationis populated once a claim withmatchAttribute/distinctAttributeconstraints over a list-typed attribute is successfully allocated; the claim staysPending(unschedulable) if no device combination satisfies the constraint.
- Other (treat as last resort)
- Details:
What are the reasonable SLOs (Service Level Objectives) for the enhancement?
Existing DRA and scheduler SLOs continue to apply; this feature does not introduce new latency-sensitive control loops. Since the added constraint evaluation (set intersection/pairwise-disjoint check) is bounded by the per-device attribute-value limit (ResourceSliceMaxAttributeValuesPerDevice = 48), it is not expected to measurably change scheduler_plugin_execution_duration_seconds{plugin="DynamicResources"} p99 relative to the scalar-only case.
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
- Metrics
- Metric names:
apiserver_request_total{resource="resourceslices"}/apiserver_request_total{resource="resourceclaims"}— non-2xx rates indicate validation/admission problems with list-typed attribute fields.scheduler_unschedulable_pods{plugin="DynamicResources"}— indicates constraint-evaluation allocation failures.scheduler_plugin_execution_duration_seconds{plugin="DynamicResources"}(extension_point="Filter") — indicates allocator performance for constraint evaluation.
- Components exposing the metric: kube-apiserver, kube-scheduler
- Metric names:
- Other (treat as last resort)
- Details:
Are there any missing metrics that would be useful to have to improve observability of this feature?
No dedicated per-feature metric (e.g., a gauge counting list-typed-attribute usage) is planned; the existing DRA/scheduler metrics above are considered sufficient signal, consistent with other DRA scheduling features (e.g. KEP-5075).
Dependencies
Does this feature depend on any specific services running in the cluster?
This feature depends on DRA structured parameters (KEP-4381) being enabled (it is GA as of v1.34), and on DRA drivers publishing list-typed DeviceAttribute values for it to be observable. It does not depend on any additional cluster-level service or node-level agent beyond the existing DRA dependencies (kube-apiserver, kube-scheduler, and a DRA driver implementing the resourceslice publishing library).
Scalability
Will enabling / using this feature result in any new API calls?
No
Will enabling / using this feature result in introducing new API types?
No
Will enabling / using this feature result in any new calls to the cloud provider?
No
Will enabling / using this feature result in increasing size or count of the existing API objects?
Yes and no. It does add new fields, which increase the worst case size of the ResourceSlice object. However, the increase is bounded, and the worst case actually gets smaller:
- Per device, the total number of attribute values is bounded by
ResourceSliceMaxAttributeValuesPerDevice(=48) instead of 32, i.e. at most 16 more values. List elements share a single key, so the key bytes do NOT multiply. - Per slice, a slice using this feature is limited to
ResourceSliceMaxDevicesWithAdvancedFeatures(=64) devices instead of 128.
ResourceClaim is NOT affected. This KEP adds no field to it, it just extends the semantics of the existing constraints[].{matchAttribute,distinctAttribute}.
Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?
Not expected. All the proposed constraints in this KEP are monotonic constraint. Thus, worst case of computational complexity for device search is the same.
Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?
No.
Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?
No.
Troubleshooting
How does this feature react if the API server and/or etcd is unavailable?
This feature adds no additional interaction with etcd beyond the existing ResourceSlice/ResourceClaim storage paths. If the API server or etcd is unavailable, this feature behaves the same as core DRA: no ResourceSlice/ResourceClaim reads/writes succeed, and kube-scheduler cannot allocate new claims (existing running workloads are unaffected). See also the general DRA troubleshooting guidance in KEP-4381, which still applies.
What are other known failure modes?
- A
ResourceClaim’smatchAttribute/distinctAttributeconstraint references a list-typed attribute, but the allocation cannot be satisfied- Detection: The claim’s pod stays
Pending/unschedulable;scheduler_unschedulable_pods{plugin="DynamicResources"}increases.kubectl describe podshows a scheduling failure event from theDynamicResourcesplugin. - Mitigations: Verify the intended device set actually has a non-empty intersection (for
matchAttribute) or is pairwise disjoint (fordistinctAttribute) by inspecting the relevantResourceSliceattribute values (kubectl get resourceslice -o yaml). - Diagnostics: kube-scheduler logs at
-v=7show per-device constraint evaluation (Allocating one device, similar to existing DRA allocator logging). - Testing: Covered by allocator unit tests (see Unit tests); e2e coverage planned before v1.38 freeze (see e2e tests).
- Detection: The claim’s pod stays
- A driver publishes an attribute that changed from scalar to list-typed, breaking an existing CEL device selector
- Detection:
ResourceClaim/DeviceClassCEL selector compilation or evaluation errors surfaced via API validation errors or scheduler logs. - Mitigations: Rewrite the CEL expression to use
.includes(...)instead of direct equality, per API Changes. - Diagnostics: kube-apiserver validation error messages on write; kube-scheduler logs for evaluation-time errors.
- Testing: Covered by CEL compiler unit tests (
compile_test.go).
- Detection:
- Version skew: an n-1 component (gate disabled) encounters a stored CEL expression referencing a list-typed attribute
- Detection: Would previously have surfaced as an evaluation error on the older/gate-disabled component; now avoided by design (see Version Skew Strategy).
- Mitigations: N/A during normal operation; if encountered, complete the rollout of the gate across all control-plane instances.
- Diagnostics: kube-apiserver/kube-scheduler logs would show CEL evaluation errors if this guarantee were violated.
- Testing: Covered by CEL compiler unit tests (
compile_test.go), which run in theStoredExpressionsenvironment with the feature disabled.
What steps should be taken if SLOs are not being met to determine the problem?
Check scheduler_plugin_execution_duration_seconds{plugin="DynamicResources"} to confirm whether allocator latency (rather than some unrelated scheduling bottleneck) is the source of the regression, then inspect the size of the ResourceSlice attribute lists and the number of matchAttribute/distinctAttribute constraints involved in the slow requests — both are bounded (ResourceSliceMaxAttributeValuesPerDevice = 48) but a cluster using values near that bound combined with many constraints is the most likely source of elevated latency.
Implementation History
- 2025-11-14: KEP created.
- Kubernetes 1.36: Alpha implementation merged behind the
DRAListTypeAttributesfeature gate. - Kubernetes 1.37:
.includesmigrated from the DRA-specific CEL library into the shared, versioned CEL lists library; list-typed attributes made evaluable in the StoredExpressions CEL environment regardless of the feature gate state; scheduler_perf coverage added for list-typedmatchAttribute. - Kubernetes 1.38: promotion to Beta.
Drawbacks
- It adds complexity to the scheduler: the allocator’s constraint evaluation and the CEL type system both have to handle an attribute that may be either a scalar or a list.
- The type of an attribute is no longer determined by its name alone. The same name can be a scalar on one driver’s device and a list on another’s, so a device selector has to use
.includesrather than==, and existing selectors break when a driver switches an attribute from scalar to list.
Alternatives
Just support formatted string list instead of introducing list type
We could add pseudo list type support only for string type attribute (e.g. comma separated string).
- Pros:
- Simple, no change in
DeviceAttribute
- Simple, no change in
- Cons:
- String list only (Can’t support list of int/version).
- prone to mis-formatted string
- extra parsing computation
Introduce matchSemantics/distinctSemantics field for flexible/declarative match
Introduce matchSemantics/distinctSemantics fields into constraints field like this:
matchSemantics field
kind: ResourceClaim
spec:
constraints:
- requests: [ "device1", "device2", "device3" ]
matchAttribute: "resource.kubernetes.io/pcieRoot"
# [NEW]
# An optional field that defines customized "match" semantics over attribute values.
# This field must not set when "distinctAttribute" is set
matchSemantics:
# mode specifies the "match" semantics
# Identical (∀i,j, v_i = v_j):
# All the attribute values among candidate devices are identical,
# supporting both list-order-sensitive and set-equivalence comparisons via `listMode`.
# NonEmptyIntersection (|∩ v_i| >= k (>=1)):
# The intersection (as a set) of list values among candidate devices is non-empty.
# The required intersection size could be configurable via `minSize`.
# For future possible cases:
# - CommonPrefix/Suffix with customizable length
# - Identical for aggregated values of the list items (min/max/sum/length)
mode: Identical | NonEmptyIntersection
options:
nonEmptyIntersection:
# if true, implicit cast from scalar to list will be performed. The default is false.
coerceScalarToList: true | false
# minSize specifies the minimum size of the intersection to evaluate as true.
# Default is 1. The value must be positive integer.
minSize: 1
identical:
coerceScalarToList: true | false # common option
# listMode specified the equality as a set(order/duplicates are ignored) or list (order significant). Default is List
listMode: List | Set
Examples of match semantics mode:
| attribute values | Identical | NonEmptyIntersection( coerceScalarToList=true) |
|---|---|---|
d1="a", d2="b" | false | false |
d1=["a", "b"] , d2=["b", "a"] | false(listMode: List)true(listMode: Set) | true( d1 ∩ d2 = {"a", "b"}) |
d1=["a", "b"] , d2=["a", "c"] | false | true( d1 ∩ d2 = {"a"}) |
d1=["a", "b"] , d1=["c", "d"] | false | false( d1 ∩ d2 = ∅) |
`distinctSemantics
kind: ResourceClaim
spec:
constraints:
- requests: [ "device1", "device2", "device3" ]
distinctAttribute: "resource.kubernetes.io/numaNode" # note: this is imaginary attribute.
# [NEW]
# an optional field that defines customized "distinct" semantics over attribute values
# this field must not set when "matchAttribute" is set
distinctSemantics:
# mode specifies the "distinct" semantics
# `AllDistinct`:
# All the values are distinct, supporting both list-order-sensitive and set-equivalence comparisons via `listMode`.
# (i.e. ∀i,j s.t. i ≠ j, v_i != v_j),
# `EmptyIntersection`:
# The intersection (as a set) of all the list values among candidate devices is empty. (i.e. ∩ v_k = ∅ )
# `PairwiseDisjoint`:
# Every pair of the list values (as a set) of candidate devices is disjoint (i.e. completely no overlap).
# (i.e. ∀i,j s.t. i ≠ j, v_i ∩ v_j = ∅),
# For future possible cases:
# - NoCommonPrefix/Suffix, PairwiseDisjointPrefix/Suffix with customizable length
# - AllDistinct for aggregated values of the list items (min/max/sum/length)
mode: AllDistinct | EmptyIntersection | PairwiseDisjoint
options:
allDistinct:
coerceScalarToList: true | false # common option
# listMode specified the equality as a set(order/duplicates are ignored) or list (order significant). Default is List
listMode: List | Set
emptyIntersection:
coerceScalarToList: true | false # common option
pairwiseDisjoint:
coerceScalarToList: true | false # common option
Examples of distinct semantics mode:
| attribute values | AllDistinct | PairwiseDistinct( coerceScalarToList=true) | EmptyIntersection( coerceScalarToList=true) |
|---|---|---|---|
d1="a", d2="b" | false | false | false |
d1=["a", "b"] , d2=["b", "a"] | true(listMode: List)false(listMode: Set) | false( d1 ∩ d2={"a","b"}) | false( ∩dk={"a","b"}) |
d1=["a", "b"] , d2=["a", "c"], d3=["a", "d"] | true | false( di ∩ dj = {"a"} ≠ ∅) | false( ∩ dk = {"a"} ≠ ∅) |
d1=["a", "b"] , d2=["b", "c"], d3=["c", "a"] | true | false( di ∩ dj ≠ ∅) | true( ∩ dk = ∅) |
d1=["a", "b"] , d2=["c", "d"], d3=["e", "f"] | true | true( di ∩ dj = ∅) | true( ∩ dk = ∅) |
Pros/Cons
- Pros:
- Flexible
- Declarative
- Extensible
- Cons:
- Too much complex even we don’t have use-cases to introduce the complexity
Unified semantics field instead of matchSemantics/distinctSemantics
We can consider unified semantics field for both matchAttribute/distinctAttribute like below:
semantics:
mode: NonEmptyIntersection | EmptyIntersection | Identical | AllDistinct | PairwiseDisjoint
- Pros:
- Simple
- Cons:
- Confusing which mode is valid for
matchAttributeordistinctAttribute - Extra validation logics
- Confusing which mode is valid for