KEP-6011: CSI ControllerGetNodeInfo
KEP-6011: CSI ControllerGetNodeInfo
- Release Signoff Checklist
- Summary
- Motivation
- Proposal
- Design Details
- Production Readiness Review Questionnaire
- Implementation History
- Drawbacks
- Alternatives
- Alternative 1: Private CRD and Controller
- Alternative 2: Instance Metadata Enhancement
- Alternative 3: Node Label Patching (e.g., AWS metadata-labeler)
- Alternative 4: CRD-based Topology Retrieval (e.g., vSphere CSINodeTopology)
- Alternative 5: Hardcoded Instance-Type Tables in the Driver Binary
- Alternative 6: Separate NodeGetID RPC
- Alternative 7: Combine Node and Controller Values
- Alternative 8: Reactive-Only Discovery
- Alternative 9: Static Node Context Object
- Alternative 10: Controller Results in CSINode Status, Published by Kubelet
- Infrastructure Needed
Release Signoff Checklist
Items marked with (R) are required prior to targeting to a milestone / release.
- (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
- (R) KEP approvers have approved the KEP status as
implementable - (R) Design details are appropriately documented
- (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input
- e2e Tests for all Beta API Operations
- (R) Ensure GA e2e tests meet requirements for [Conformance Tests]
- (R) Minimum Two Week Window for GA e2e tests to prove flake free
- (R) Graduation criteria is in place
- (R) [all GA Endpoints] must be hit by [Conformance Tests] within one minor version of promotion to GA
- (R) Production readiness review completed
- (R) Production readiness review approved
- “Implementation History” section is up-to-date for milestone
- User-facing documentation has been created in kubernetes/website
- Supporting documentation—e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes
Summary
This KEP introduces a new optional CSI RPC, ControllerGetNodeInfo, and a companion request flag on the existing NodeGetInfo, that together allow a CSI driver to split node registration into two phases: a lightweight node-side identity call and a controller-side lookup for topology and volume attachment limits. This eliminates the need for cloud API credentials on worker nodes while preserving full topology-aware scheduling and accurate volume limit tracking.
Motivation
Today, the CSI NodeGetInfo RPC is the single entry point for a node to report its identity, topology, and volume attachment limit to the Container Orchestrator (CO). In practice, some CSI driver implementations require cloud API credentials on the node to fully populate this response, for example to query the instance’s availability zone or the maximum number of attachable volumes. Other implementations use local instance metadata, static reservations, or provider-specific controllers.
This creates several problems:
Security: Organizations with strict security postures, particularly in financial services and government, prohibit distributing cloud API credentials to worker nodes. These users must choose between security and full CSI functionality.
Non-CSI attachment accounting: Out-of-band attached volumes are not counted by the scheduler, so the reported limit must exclude them. SP sees every attachment through the cloud API. But it maybe unable to distinguish Out-of-band attached volumes from CO-published ones Only the CO knows which ones it published. They need to combine the information to calculate the volume limit for scheduler.
Scalability: In large clusters, every node independently calls cloud APIs during registration. A 5000-node cluster startup produces 5000 concurrent API calls, risking throttling and slow registration. A controller-side approach enables batching, caching, and coordinated rate limiting.
Dynamic updates: KEP-4876 made
CSINode.Spec.Drivers[*].Allocatable.Countmutable and introduced periodic and failure-triggered updates viaNodeGetInfo. This KEP moves the update source to the controller side while preserving those refresh mechanisms.
This KEP addresses all four problems by introducing a clean split: the node reports only its identity (cheap, local, no credentials), and the controller fills in topology and volume attachment limits (where credentials and VolumeAttachment context already exist).
Goals
- Enable CSI node registration without cloud API credentials on the node
- Support non-CSI attachment accounting using controller-side
VolumeAttachmentknowledge - Improve scalability through controller-side batching and caching of cloud API calls
- Maintain full backward compatibility; drivers that do not adopt the new flow continue to work unchanged
Non-Goals
- Modifying Kubernetes core scheduling logic
- Requiring changes to CSI drivers that do not need this feature
- Implementing cloud provider-specific solutions within Kubernetes core
- Using
ControllerGetNodeInfoin Kubernetes for deployments withCSIDriver.Spec.AttachRequired=false
Proposal
User Stories
Story 1: Security-hardened Environment
A financial services company prohibits cloud API credentials on worker nodes. Today, their CSI driver’s NodeGetInfo either returns incomplete data (no topology, inaccurate limits) or requires a provider-specific workaround. With this proposal, the node answers NodeGetInfo with only its node_id from local instance metadata, and the controller calls ControllerGetNodeInfo with existing credentials. Full functionality, no credentials on nodes, no custom workarounds.
Story 2: Accurate Non-CSI Volume Accounting
An operator provisions nodes with boot volumes and additional data disks outside Kubernetes.
Today, operators rely on static reservations or provider-specific mechanisms to account for these attachments.
With this proposal, the SP returns the volumes attached to the node, and external-attacher subtracts those without a VolumeAttachment from the SP’s limit.
Story 3: Large Cluster Scalability
A 5000-node cluster startup triggers 5000 concurrent cloud API calls from NodeGetInfo. With this proposal, the controller can batch instance queries, cache results by instance type, and apply coordinated rate limiting, significantly reducing cloud API load and registration time.
Risks and Mitigations
| Risk | Mitigation |
|---|---|
| Node registration latency increases | The node-side NodeGetInfo becomes faster (no cloud API). The controller-side roundtrip adds seconds, but node registration is a one-time event. Net impact is minimal. |
| Controller becomes a bottleneck | The controller already handles ControllerPublishVolume for every attach. ControllerGetNodeInfo adds one call per node registration, which is negligible overhead. Batching and caching further reduce load. |
Race between ControllerGetNodeInfo and concurrent attach/detach | CO records volume IDs processed during the call and considers them CSI-managed. See Race Condition Mitigation . |
| Version skew | Feature gates on kube-apiserver, kubelet, and external-attacher. See Version Skew Strategy . |
Notes/Constraints/Caveats
CSI Spec Dependency: This KEP requires CSI spec PR #603 to be merged first. Kubernetes implementation cannot proceed until the spec changes land.
Upgrade Order: External-attacher must be upgraded before nodes. If external-attacher does not yet support
ControllerGetNodeInfo, registrations without an existingSpec.Driversentry remain pending. Existing entries with a matching node ID are preserved but cannot be refreshed through the controller-side flow until external-attacher catches up.Backward Compatibility: Drivers that do not adopt the new flow, and deployments with
AttachRequired=false, continue usingNodeGetInfounchanged. No breaking changes to the existing flow.Kubernetes Scope: This workflow requires VolumeAttachments and therefore applies only to deployments where
CSIDriver.Spec.AttachRequiredis not false (the default is true). UsingControllerGetNodeInfowithAttachRequired=falseis outside the scope of this KEP. We also requires CSIPUBLISH_UNPUBLISH_VOLUMEcapability, but only because lacking interesting use-case withoutPUBLISH_UNPUBLISH_VOLUME. This also reduces the testing burden. We should reconsider if a use-case appears.Node-local Limit Overrides: Drivers are encouraged to use
published_volume_idsto replace manual reservations for non-CSI attachments where possible. The new flow does not forward node-side overrides; settings still required need equivalent controller-side configuration, supplied through driver-provided deployment templates (e.g., Helm) or manually by the administrator. Deployments whose required per-node behavior cannot be reproduced on the controller should retain theNodeGetInfoflow.
Design Details
CSI Spec Changes
This KEP depends on CSI spec PR #603 .
NodeGetInfo Request Flag
A new optional field is added to NodeGetInfoRequest:
message NodeGetInfoRequest {
// When true, the CO will obtain accessible_topology and
// max_volumes_per_node from ControllerGetNodeInfo. The SP MAY omit
// those two fields from NodeGetInfoResponse and return only node_id.
// The CO MUST NOT consume accessible_topology or max_volumes_per_node
// from a response to a request with this field set.
// The CO MUST NOT set this field to true unless the SP has the
// NODE_INFO_FROM_CONTROLLER node capability.
// This field is OPTIONAL.
bool controller_get_node_info = 1 [(alpha_field) = true];
}
The CO sets controller_get_node_info only when the driver advertises the NODE_INFO_FROM_CONTROLLER node capability (see New Capabilities
).
The node_id returned in this mode comes from local instance metadata and requires no cloud API credentials — for example, http://169.254.169.254/latest/meta-data/instance-id (AWS) or http://100.100.100.200/latest/meta-data/instance-id (Alibaba Cloud).
Why a request flag rather than always deferring: the flag is a handshake. Under the new workflow an SP may want to skip fetching topology and limits itself, either to avoid storing credentials on the node or to reduce API traffic to the cloud or Kubernetes. Without the flag, such an SP would have to either:
- Omit topology and limits unconditionally, assuming every CO supports the new workflow.
But
NodeGetInfoResponsehas no safe “I don’t know yet” value for these fields —max_volumes_per_node = 0means “no limit” and emptyaccessible_topologymeans “no constraint” — so an old CO would silently mis-schedule. - Always return them, incurring useless API traffic even where the CO ignores the result.
The flag resolves this: the SP omits those fields only when the CO has signaled it will
source them authoritatively from ControllerGetNodeInfo, so the node side is never misread.
Why not a separate NodeGetID RPC: reusing NodeGetInfo requires only a capability plus one optional request field,
so the existing CSI drivers add support incrementally instead of implementing a new RPC.
node_id also flows to the controller side in this design (see below); a shared node-produced field has one natural home on NodeGetInfoResponse, whereas a subset RPC would duplicate it.
See Alternative 6
.
ControllerGetNodeInfo RPC
rpc ControllerGetNodeInfo(ControllerGetNodeInfoRequest) returns (ControllerGetNodeInfoResponse) {
option (alpha_method) = true;
}
message ControllerGetNodeInfoRequest {
// Node ID returned by NodeGetInfo.
string node_id = 1;
}
// All response fields are optional.
message ControllerGetNodeInfoResponse {
// Volume attachment limit calculated by the SP; zero leaves it unspecified.
int64 max_volumes_per_node = 1;
// Accessible topology (zone, region, etc.).
Topology accessible_topology = 2;
// Volumes attached according to the cloud API.
repeated string published_volume_ids = 3;
}
Retrieves topology, attached volumes, and instance limit from the controller side, where cloud API credentials are already available.
Design: volume classification. The scheduler treats allocatable.count in CSINode as the number of CSI-managed volumes the node can support, then subtracts the CSI volumes it already knows about to determine available slots. The scheduler has no awareness of non-CSI attachments (boot volumes, network interfaces consuming shared device slots, manually attached disks). So non-CSI volumes must be accounted for.
Existing drivers account for non-CSI attachments using node-local information, static reservations, or provider-specific controllers. This KEP standardizes how SP-reported attachments are combined with the CO’s volume records.
The SP calculates the volume attachment limit (accounting for ENIs etc.) and reports the attached volumes. The CO (external-attacher) has the VolumeAttachment context needed to identify non-CSI volumes from that list and subtract them:
volume_limit = max_volumes_per_node (SP calculated, accounting for ENIs etc.)
total_attached = published_volume_ids (from SP response)
csi_managed = VolumeAttachment objects (CO knows)
non_csi_attached = total_attached - intersection(total_attached, csi_managed)
effective_limit = volume_limit - non_csi_attached
External-attacher writes effective_limit to CSINode.Spec.Drivers[*].Allocatable.Count; the scheduler’s volume counting is unchanged. published_volume_ids is optional: an SP that accounts for non-CSI volumes itself can omit the list, in which case external-attacher uses max_volumes_per_node unchanged.
Example: Instance type limit is 25. Node has 2 ENIs (consuming 2 slots on shared-limit types). SP calculates attachment limit = 23. Cloud API shows 10 attached volumes (published_volume_ids). CO has 8 CSI volumes in VolumeAttachment. CO identifies 2 non-CSI volumes (boot volume + manually attached disk) → effective limit = 23 - 2 = 21. Scheduler subtracts 8 CSI volumes → 13 available. Correct: 25 - 2 (ENIs) - 10 (attached) = 13 real remaining.
CO avoids a race condition by recording all volume IDs processed during the ControllerGetNodeInfo call and considers them CSI-managed.
Example implementations:
- AWS EBS:
DescribeInstancesfor AZ/region and current block device mappings,DescribeInstanceTypesfor attachment limit. SP calculates volume attachment limit accounting for ENI-consumed slots on shared-limit instance types, returns this limit and all attached volume IDs. This would replace the existing metadata-labeler sidecar and the--reserved-volume-attachmentsCLI flag. - Alibaba Cloud:
DescribeInstancesfor zone/region,DescribeAvailableResourcefor disk categories,DescribeDisksfor current attachments,DescribeInstanceTypesfor limits.
New Capabilities
NodeServiceCapability.RPC.NODE_INFO_FROM_CONTROLLER: the node plugin supports thecontroller_get_node_inforequest flag and may omit topology/limits when it is set. kubelet needs this node-side signal because it cannot observe controller capabilities.ControllerServiceCapability.RPC.GET_NODE_INFO: indicates support forControllerGetNodeInfo
Invariant: If a driver advertises NODE_INFO_FROM_CONTROLLER, it MUST also advertise GET_NODE_INFO. This is enforced by the CSI spec. Without this invariant, a node could register with only a node_id in CSINode.Spec.DriverRegistrations and never have topology or allocatable populated, because the SP omits them once kubelet sets the flag.
If a driver advertises GET_NODE_INFO, it MUST also advertise PUBLISH_UNPUBLISH_VOLUME.
This is not a hard constraint from API, but a lack of interesting use-cases; can be reconsidered if a use-case appears.
Topology key consistency: ControllerGetNodeInfo MUST return the same topology keys as the driver’s NodeGetInfo would. Existing PersistentVolumes have nodeAffinity rules referencing these keys (e.g., topology.ebs.csi.aws.com/zone). Inconsistent keys would break scheduling for already-provisioned volumes.
Kubernetes Integration
kubelet Changes
When the CSIControllerGetNodeInfo feature gate is enabled, the CSI node plugin advertises NODE_INFO_FROM_CONTROLLER, and CSIDriver.Spec.AttachRequired is not false:
- Call
NodeGetInfowithcontroller_get_node_info = true - Store the
node_idinCSINode.Spec.DriverRegistrations(see CSINode Driver Registrations ). Verify that the API response contains the input; if the API server dropped the field, fail registration - Do NOT populate topology or allocatable from the response, even if present; external-attacher handles this via
ControllerGetNodeInfo - Skip the KEP-4876
NodeGetInfocalls (periodic and afterRESOURCE_EXHAUSTED) for this driver, as external-attacher takes over. kubelet retains the existing KEP-4876 Pod failure behavior forRESOURCE_EXHAUSTED. - If
NodeGetInfofails or returns an emptynode_id, fail registration - If
NODE_INFO_FROM_CONTROLLERis not advertised orAttachRequired=false, use the existingNodeGetInfoflow unchanged
req := &csi.NodeGetInfoRequest{}
if hasNodeInfoFromControllerCapability(driver) && attachRequired(driver) {
req.ControllerGetNodeInfo = true
}
info, err := nodePlugin.NodeGetInfo(req)
if err != nil {
return fmt.Errorf("NodeGetInfo failed: %w", err)
}
if req.ControllerGetNodeInfo {
// topology/allocatable are ignored; external-attacher supplies them.
// Record the node ID in CSINode.Spec.DriverRegistrations.
return setRegistrationNodeID(driverName, info.NodeId)
} else {
// ... existing flow: populate topology and allocatable from info ...
}
When entering the controller-side flow, kubelet preserves the existing Spec.Drivers entry while writing the registration input. Enabling the feature or restarting kubelet with the same node ID does not invalidate node-populated information. If the write fails, including validation rejection because the node IDs differ, the error propagates through the normal registration-failure path: kubelet unregisters the driver, removing both entries, and reports registration failure to the registrar. Registration is then retried, consistent with the traditional flow.
When the driver unregisters, kubelet removes both its DriverRegistrations entry and its Spec.Drivers entry in one CSINode update. When kubelet registers a driver through the existing flow, including after the feature gate is disabled, it removes the driver’s DriverRegistrations entry in the same update that writes Spec.Drivers, which stops external-attacher from processing the driver on this node.
external-attacher Changes
When the CSIControllerGetNodeInfo feature gate is enabled, CSINode.Spec.DriverRegistrations has an entry for the driver, and the CSI controller plugin advertises GET_NODE_INFO:
- Call
ControllerGetNodeInfowhen the driver has aDriverRegistrationsentry but noCSINode.Spec.Driversentry (initial registration) - Call
ControllerGetNodeInfoafterControllerPublishVolumereturnsRESOURCE_EXHAUSTEDifCSIDriver.Spec.NodeAllocatableUpdatePeriodSecondsis set (capacity correction, building on KEP-4876) - Call
ControllerGetNodeInfoperiodically ifCSIDriver.Spec.NodeAllocatableUpdatePeriodSecondsis set (periodic refresh, building on KEP-4876) - Calculate effective
max_volumes_per_nodeby comparingpublished_volume_idsfrom SP response againstVolumeAttachmentobjects - For initial registration, write the topology values as Node labels, preserving the existing kubelet topology collision checks (Needs new RBAC permission)
- Create the
CSINode.Spec.Driversentry with the topology keys and calculated capacity, or update existing entry’s capacity (Needs new RBAC permission)
If the driver has a DriverRegistrations, report events on CSINode when:
ControllerGetNodeInfofailed;- Driver does not advertise
GET_NODE_INFO; - Driver advertise
GET_NODE_INFObut notPUBLISH_UNPUBLISH_VOLUME.
The CSINode update uses the resourceVersion from the snapshot passed to ControllerGetNodeInfo.
On a CSINode update conflict, discard the RPC result and wait for next CSINode update event from informer.
Reconciliation starts again from the latest CSINode and calls ControllerGetNodeInfo again if still needed.
This prevents in-flight results from restoring an unregistered driver or overwriting a switch to the traditional flow. Node and CSINode writes are not atomic: a successful label patch can remain after a failed CSINode update.
If this driver has no completed entry in spec.drivers, retry non-conflict errors with exponential backoff.
Otherwise, retain the previously published information and retry on the next periodic or reactive trigger.
The registration input remains present after completion and supplies node_id for subsequent ControllerGetNodeInfo calls. ControllerPublishVolume continues to use the completed Spec.Drivers entry.
Adding a registration input or restarting external-attacher does not force a lookup for an existing completed entry. On startup, external-attacher processes pending registrations, while completed entries follow the configured periodic and RESOURCE_EXHAUSTED refresh triggers. Preserved node-populated values can remain indefinitely if neither trigger occurs. An unregistration/re-registration cycle that removes the completed entry also triggers discovery; a kubelet restart alone does not.
type nodeInfoProcessor struct {
//...
}
func (p *nodeInfoProcessor) processNode(csiNode *CSINode, forceRefresh bool) {
reg, ok := findRegistration(csiNode.Spec.DriverRegistrations, driverName)
if !ok {
return
}
nodeID := reg.NodeID
if driverInSpec(csiNode) && !periodicUpdateDue() && !forceRefresh {
return
}
publishing, stop := p.startRecording(nodeName)
csiPublished := listVolumeAttachments(nodeName)
info := ControllerGetNodeInfo(nodeID)
stop()
var effectiveLimit *int32
if info.maxVolumesPerNode != 0 {
// Calculate effective limit:
// SP already accounted for ENIs etc. in maxVolumesPerNode.
// CO subtracts non-CSI volumes (attached but not in VolumeAttachment).
csiPublished := csiPublished.Union(publishing)
nonCsi := info.publishedVolumeIDs.Difference(csiPublished)
effectiveLimit = new(max(0, info.maxVolumesPerNode - len(nonCsi)))
}
if driverInSpec(csiNode) {
updateAllocatable(csiNode, effectiveLimit)
} else {
// Node labels first, then CSINode.Spec.Drivers with csiNode's resourceVersion.
updateNodeLabels(nodeName, info.accessibleTopology)
updateCSINode(csiNode, nodeID, info.accessibleTopology, effectiveLimit)
}
}
Integration with KEP-4876: When a driver supports ControllerGetNodeInfo, external-attacher takes over the responsibilities that KEP-4876 assigns to kubelet:
| Responsibility | KEP-4876 (kubelet) | KEP-6011 (external-attacher) |
|---|---|---|
| Periodic updates | NodeGetInfo at NodeAllocatableUpdatePeriodSeconds interval | ControllerGetNodeInfo at same interval |
| RESOURCE_EXHAUSTED handling | kubelet detects error, calls NodeGetInfo | external-attacher detects error, calls ControllerGetNodeInfo |
When NodeAllocatableUpdatePeriodSeconds is set, external-attacher attempts to refresh the node’s allocatable count before reporting RESOURCE_EXHAUSTED through the VolumeAttachment error code. This reactive refresh bypasses the periodic-update check (forceRefresh = true). Refresh failure does not suppress the attach error, preserving existing best-effort recovery behavior.
The key advantage: external-attacher has accurate VolumeAttachment context, enabling precise non-CSI volume classification and accurate capacity calculation.
Periodic update scalability: External-attacher uses a rate-limited work queue with jitter (±20% of the configured period) rather than per-node timers. On external-attacher startup, pending registrations are enqueued immediately. For completed registrations with periodic updates enabled, the first periodic refresh is scheduled at a uniformly random offset within the configured period. This prevents thundering herd on restart and provides natural rate limiting for cloud API calls.
Race Condition Mitigation
A race exists between ControllerGetNodeInfo and concurrent attach/detach: if an attach completes between listing VolumeAttachment objects and the cloud API query, the newly attached volume appears in SP’s published_volume_ids but not in the CO’s CSI records, causing the CO to misclassify it as non-CSI.
Mitigation: The CO records all volume IDs processed during the ControllerGetNodeInfo call and considers them CSI-managed.
When classifying volumes, the CO considers any volume ID processed during the ControllerGetNodeInfo call as CSI-managed.
For CO-managed attachments, this approach covers:
- volumes that have
VolumeAttachmentbefore the call, including those with uncertain status (in-progress or failed attaches) - volumes attached during the call,
- volumes detached during the call,
- and even volumes that were attached then detached during the call.
These volumes are classified as CSI-managed, assuming the SP does not return successfully unpublished volumes in subsequent ControllerGetNodeInfo calls.
Keeping already detached CSI volume IDs in the set does not affect the non-CSI volume count.
This does not prevent out-of-band attachments after the cloud API snapshot. The reported limit may therefore be temporarily stale; periodic and failure-triggered updates from KEP-4876, when configured, allow it to be refreshed.
Workflow
Initial registration on new node:
sequenceDiagram
box rgba(255,0,0,0.1) Node Side
participant kubelet
participant csi-node as CSI Node Plugin
end
box Controller Side
participant apiserver as API Server
participant attacher as external-attacher
participant csi-ctrl as CSI Controller Plugin
end
kubelet->>+csi-node: NodeGetCapabilities
csi-node-->>-kubelet: NODE_INFO_FROM_CONTROLLER capability
attacher->>+csi-ctrl: ControllerGetCapabilities
csi-ctrl-->>-attacher: GET_NODE_INFO capability
kubelet->>+csi-node: NodeGetInfo(controller_get_node_info=true)
csi-node->>csi-node: Read instance ID from metadata
csi-node-->>-kubelet: node_id (topology/limits omitted)
kubelet->>apiserver: Update CSINode spec.driverRegistrations
apiserver-->>attacher: CSINode watch event
attacher->>attacher: List VolumeAttachments for node
attacher->>+csi-ctrl: ControllerGetNodeInfo(node_id)
csi-ctrl->>csi-ctrl: Query cloud APIs
csi-ctrl-->>-attacher: topology, max_volumes_per_node, published_volume_ids
attacher->>attacher: Calculate effective limit (subtract non-CSI from SP limit)
attacher->>apiserver: Update Node labels and CSINode.Spec.Drivers
Note over attacher,csi-ctrl: RESOURCE_EXHAUSTED scenario
attacher->>+csi-ctrl: ControllerPublishVolume
csi-ctrl-->>-attacher: RESOURCE_EXHAUSTED
attacher->>+csi-ctrl: ControllerGetNodeInfo(node_id)
csi-ctrl-->>-attacher: Updated max_volumes_per_node, published_volume_ids
attacher->>apiserver: Update CSINode Allocatable
attacher->>apiserver: Update VA.status.attachError.errorCode
apiserver-->>kubelet:
kubelet->>apiserver: fail relevant podsAPI Changes
CSINode Driver Registrations
This KEP adds a feature-gated driverRegistrations list to CSINodeSpec, which carries node_id from kubelet to external-attacher. kubelet adds an entry after it calls NodeGetInfo with the flag, and removes it when the driver unregisters or registers through the existing flow:
type CSINodeSpec struct {
// ... existing fields ...
// driverRegistrations contains node-side inputs for controller-side discovery.
// This field is alpha-level and requires the CSIControllerGetNodeInfo feature gate.
// +featureGate=CSIControllerGetNodeInfo
// +optional
// +patchMergeKey=name
// +patchStrategy=merge
// +listType=map
// +listMapKey=name
DriverRegistrations []CSINodeDriverRegistration `json:"driverRegistrations,omitempty" patchStrategy:"merge" patchMergeKey:"name"`
}
type CSINodeDriverRegistration struct {
Name string `json:"name"`
NodeID string `json:"nodeID"`
}
apiVersion: storage.k8s.io/v1
kind: CSINode
metadata:
name: my-node
spec:
driverRegistrations: # added and removed by kubelet
- name: csi.example.com
nodeID: i-instanceid
drivers:
- name: csi.example.com # added by external-attacher, removed by kubelet
nodeID: i-instanceid
topologyKeys:
- topology.kubernetes.io/zone
allocatable:
count: 6
The two lists together describe the registration state of a driver on the node:
drivers entry | driverRegistrations entry | State |
|---|---|---|
| absent | absent | The plugin is not registered, and kubelet starts one of the two flows. |
| present | absent | The existing node-side flow is in use, and external-attacher does nothing. |
| absent | present | NodeGetInfo is done and ControllerGetNodeInfo is pending, so external-attacher calls it. |
| present | present | Driver information is available; it may still be the preserved node-side result until a controller refresh occurs. |
Validation and feature gating follow the Kubernetes new-field API guidance :
- The list is optional, with no default. Each entry requires a valid CSI driver name and node ID, using the same validation as
CSINodeDriver; duplicate names are rejected. - When both lists contain an entry for a driver, their node IDs must match. Updates introducing a mismatch are rejected. Existing completed-entry immutability rules remain unchanged.
- With the API-server gate disabled, drop the field on create, and on update only if the old object did not already use it. Existing usage can still be updated or removed.
- Validate the field whenever present, independently of the gate. Consumers separately honor their own gates.
We should not put an incomplete entry into spec.drivers and add a new ready: false field. Old consumers would ignore that field. Missing topology keys and an unset allocatable count already have valid meanings—no topology and an unbounded volume count, not “still initializing”.
Test Plan
[X] I/we understand the owners of the involved components may require updates to existing tests.
Prerequisite testing updates
- CSI mock driver updated to support the
controller_get_node_inforequest flag and theControllerGetNodeInfoRPC
Unit tests
- API validation and storage strategy: Entry validation, duplicate names, matching node IDs, field dropping with the gate disabled, and preservation/update/removal of existing field usage after disablement
k8s.io/kubernetes/pkg/volume/csi: Capability detection,NodeGetInfowith the request flag,DriverRegistrationshandling,resourceVersionconflict retryk8s.io/kubernetes/pkg/kubelet:NodeGetInfofailure blocks registration, default flow whenNODE_INFO_FROM_CONTROLLERis absent orAttachRequired=false, periodic update responsibility switchingexternal-attacher:DriverRegistrationsdetection andControllerGetNodeInfotrigger, effective limit calculation (comparingpublished_volume_idsfrom SP against VolumeAttachments), race condition mitigation (recording processed volume IDs),RESOURCE_EXHAUSTED→ControllerGetNodeInfo→ CSINode update flow, multi-driver coexistence (one driver uses the new flow, another does not), periodic update work queue with jitter, partial response handling, external-attacher restart recovery
Integration tests
- Node registration end-to-end with the controller-side flow
- Traditional → controller-side → traditional transitions preserve matching-identity information without forcing a controller lookup, including kubelet restart and no periodic refresh configuration
- A changed node ID causes validation rejection, normal registration-failure cleanup, and successful registration on retry in both flows
- External-attacher startup processes pending registrations without forcing refreshes of completed entries
- Unregistration or rollback during an in-flight controller lookup, and failed Node patches preventing completed publication
NodeGetInfofailure or a dropped registration input blocks registration- Capacity update after
RESOURCE_EXHAUSTED - Omitting
published_volume_idspreserves the SP-reported limit without CO-side deductions
e2e tests
- End-to-end workflow with CSI driver supporting the controller-side flow
- Backward compatibility with drivers not supporting the controller-side flow
AttachRequired=falseretainsNodeGetInfowithout external-attacher even when the feature gate andNODE_INFO_FROM_CONTROLLERcapability are enabled- Topology-aware scheduling with controller-side topology
- Capacity update after volume limit reached
Graduation Criteria
Alpha
- Feature implemented behind the
CSIControllerGetNodeInfofeature gate (kube-apiserver, kubelet, and external-attacher) - CSI spec PR #603 merged (alpha)
- kubelet:
controller_get_node_inforequest flag with node-only flow when the capability is absent - external-attacher:
ControllerGetNodeInfosupport - Unit and integration tests passing
Beta
- CSI spec RPCs promoted to beta
- Feedback incorporated from at least two CSI driver implementations
- All e2e tests passing
- Scalability validated in clusters with 1000+ nodes
- CSI driver developer documentation published
GA
- CSI spec RPCs stable
- Multiple CSI drivers using the feature in production
- No critical issues for two consecutive releases
- Documentation complete in kubernetes/website
Upgrade / Downgrade Strategy
Upgrade: Control-plane-first. Enable the feature gate on kube-apiserver, upgrade external-attacher (with feature gate enabled), then upgrade nodes incrementally. The controller is ready to process DriverRegistrations entries before nodes start producing them. No coordination beyond ordering is required.
Downgrade: Reverse order. Downgrade nodes first (they revert to the node-only flow and clear the new field), then downgrade external-attacher and kube-apiserver. Existing CSINode objects remain valid throughout.
Directly downgrade from a feature-enabled deployment to a version that does not understand the new field is not supported. Disable the feature-gate first, then downgrade.
If CSI driver is re-configured after using this feature (e.g. credential removed from node), that should be reverted before downgrading.
Version Skew Strategy
| Scenario | Behavior |
|---|---|
kubelet has feature, CSI driver lacks NODE_INFO_FROM_CONTROLLER | kubelet detects missing capability, calls NodeGetInfo without the flag |
CSI driver has NODE_INFO_FROM_CONTROLLER, kubelet lacks feature | capability ignored, NodeGetInfo called without the flag; SP returns full node-side info |
external-attacher has feature, CSI controller lacks GET_NODE_INFO | external-attacher detects missing capability, skips ControllerGetNodeInfo |
CSI controller has GET_NODE_INFO, external-attacher lacks feature | GET_NODE_INFO capability ignored |
| Node side has feature, controller side does not | Registrations without completed entries remain pending; matching existing entries are preserved but not refreshed |
| API server rejects or drops the registration input | Kubelet fails registration, registrar retries |
Production Readiness Review Questionnaire
Feature Enablement and Rollback
How can this feature be enabled / disabled in a live cluster?
- Feature gate
- Feature gate name:
CSIControllerGetNodeInfo - Components depending on the feature gate: kube-apiserver, kubelet, external-attacher
- Feature gate name:
Does enabling the feature change any default behavior?
When the gate is enabled, CSI drivers advertising NODE_INFO_FROM_CONTROLLER and CSIDriver.Spec.AttachRequired=true will start to use the new controller-side flow.
Other deployments continue to use NodeGetInfo unchanged.
Can the feature be disabled once it has been enabled?
Yes. Set feature gates to false and restart components. kubelet reverts to calling NodeGetInfo without the flag. Existing CSINode objects remain valid.
What happens if we reenable the feature if it was previously rolled back?
kubelet re-checks capabilities and AttachRequired, and uses the controller-side flow if supported. External-attacher re-processes any pending DriverRegistrations entries.
Are there any tests for feature enablement/disablement?
Yes, unit tests cover capability detection, fallback logic, and behavior with feature gate on/off.
Rollout, Upgrade and Rollback Planning
How can a rollout or rollback fail? Can it impact already running workloads?
Running workloads are not affected. Failure scenarios affect only new node registrations and new scheduling decisions:
- kubelet enabled but external-attacher not upgraded: registrations without completed entries remain pending; matching existing entries are preserved but not refreshed. Mitigated by upgrading controller first.
NodeGetInfofails: node registration fails for that driver. Mitigated by fixing the driver or disabling the feature gate.
What specific metrics should inform a rollback?
csi_operations_seconds{method_name="NodeGetInfo",grpc_status_code!="OK"}: high error rate indicates node-side registration failurescsi_sidecar_operations_seconds{method_name="ControllerGetNodeInfo",grpc_status_code!="OK"}: high error rate indicates controller-side failures- Increase in pods stuck in
ContainerCreatingorPendingdue to missing topology or incorrectly scheduled.
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
Manual testing during alpha: enable feature gates → verify controller-side flow used → disable feature gates → verify node-only flow → re-enable → verify controller-side flow resumes.
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?
No.
Monitoring Requirements
How can an operator determine if the feature is in use by workloads?
Check spec.driverRegistrations on CSINode objects. If the driver is listed, it is using the new flow.
How can someone using this feature know that it is working?
- Events
- Event Reason:
CSINodeInfoUpdated, emitted by external-attacher when topology/allocatable is populated
- Event Reason:
- API fields
CSINode.Spec.DriverRegistrationspopulated.- On new nodes, or after CSI driver re-registration,
CSINode.Spec.Driversalso populated, with itsTopologyKeysandAllocatable.Countfrom controller.
What are the reasonable SLOs?
- Node registration with topology populated: < 30 seconds after kubelet starts
- Capacity correction after
RESOURCE_EXHAUSTED: < 10 seconds
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
- Metrics
csi_operations_seconds{method_name="NodeGetInfo"}: kubelet-side latency and error ratecsi_sidecar_operations_seconds{method_name="ControllerGetNodeInfo"}: external-attacher-side latency and error rate- Both are existing histogram metrics with new
method_namelabel values. Error rate derived viagrpc_status_code!="OK".
Are there any missing metrics that would be useful to have to improve observability of this feature?
No. The existing csi_operations_seconds and csi_sidecar_operations_seconds histograms provide success/failure tracking, latency, and adoption visibility through the new method_name label values.
Dependencies
Does this feature depend on any specific services running in the cluster?
- CSI drivers supporting the controller-side flow: Required for the feature to activate. Drivers without the
NODE_INFO_FROM_CONTROLLERcapability, or deployments withAttachRequired=false, use the node-only flow with no impact. - external-attacher sidecar: Must be deployed with the
CSIControllerGetNodeInfofeature gate enabled. If external-attacher is down, registrations without completed entries remain pending. Matching existing entries retain their published information but cannot be refreshed.
Scalability
Will enabling / using this feature result in any new API calls?
One ControllerGetNodeInfo gRPC call per node registration, plus additional calls on RESOURCE_EXHAUSTED and periodic updates (if configured). One CSINode update per node from kubelet to add the DriverRegistrations entry, then one Node PATCH for the topology labels and one CSINode update for Spec.Drivers from external-attacher.
Will enabling / using this feature result in introducing new API types?
No.
Will enabling / using this feature result in any new calls to cloud provider?
No net new calls. The cloud API calls move from node to controller. Controller-side batching and caching can reduce total call volume in large clusters.
Will enabling / using this feature result in increasing size or count of the existing API objects?
CSINode gains one driverRegistrations entry per driver, roughly 100-200 bytes. Spec.Drivers is unchanged in size, although external-attacher populates it instead of kubelet.
Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?
Node registration time may increase slightly due to the controller-side roundtrip, offset by the node-side NodeGetInfo no longer calling cloud APIs. Overall impact expected < 1 second.
Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?
Minimal. The node-side NodeGetInfo becomes lighter (no cloud API). External-attacher handles DriverRegistrations processing and ControllerGetNodeInfo calls, plus a small map for pending node tracking.
Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?
No. NodeGetInfo remains a single gRPC call on the existing CSI socket. No new processes, files, or connections.
Troubleshooting
How does this feature react if the API server and/or etcd is unavailable?
kubelet and external-attacher retry until available. Existing workloads are unaffected. New scheduling may be delayed.
How does this feature work if the external-attacher / CSI controller is down?
Registrations without an existing completed entry remain pending with only a DriverRegistrations entry. Matching existing Spec.Drivers entries retain their published topology and allocatable, but cannot be refreshed until the controller recovers.
Topology impact: Pods with PV nodeAffinity requiring topology labels (e.g., topology.kubernetes.io/zone) may fail to schedule if the node lacks those labels. This is expected behavior -— the scheduler cannot place pods without proper topology matching.
Allocatable impact: If Allocatable.Count is not set, the scheduler’s CSI volume limits plugin currently treats this as “no limit” by default and may schedule pods that exceed the node’s actual volume capacity.
When external-attacher recovers:
- It processes pending
DriverRegistrationsentries and callsControllerGetNodeInfoto populateAllocatable.Count - It processes pending VolumeAttachments, calls
ControllerPublishVolume - If the node’s actual capacity is exhausted (due to pods scheduled during the degraded period),
ControllerPublishVolumereturnsRESOURCE_EXHAUSTED - If
NodeAllocatableUpdatePeriodSecondsis set, kubelet marks the Pod failed; its workload controller may create a replacement, which the scheduler places subject to available capacity.
This self-correcting mechanism ensures the cluster eventually reaches a consistent state.
KEP-5030
: This KEP proposes to close the gap in the scheduler’s NodeVolumeLimits plugin, so that scheduler will not place pods on nodes which aren’t reporting CSI driver information. When enabled, the degraded state will be more graceful -— pods will simply not schedule until topology/allocatable is populated.
What are other known failure modes?
NodeGetInfoRPC fails during registration- Detection:
csi_operations_seconds{method_name="NodeGetInfo",grpc_status_code!="OK"} - Mitigation: Fix driver or disable feature gate. Node retries on restart.
- Diagnostics: kubelet error log: “NodeGetInfo failed: …”
- Detection:
external-attacher cannot reach CSI controller
- Detection:
csi_sidecar_operations_seconds{method_name="ControllerGetNodeInfo",grpc_status_code!="OK"} - Mitigation: Fix the underlying issue, then restart external-attacher. Pending
DriverRegistrationsentries are re-processed on recovery. - Diagnostics: external-attacher error logs with RPC failure details.
- Detection:
Cloud API throttling during large cluster startup
- Detection: High latency in
csi_sidecar_operations_seconds{method_name="ControllerGetNodeInfo"} - Mitigation: CSI driver implements retry with backoff. Increase cloud API quotas if needed.
- Detection: High latency in
What steps should be taken if SLOs are not being met to determine the problem?
- Check error rate metrics for
NodeGetInfo(kubelet) andControllerGetNodeInfo(external-attacher) - Check latency metrics for both RPCs
- Review kubelet and external-attacher logs for RPC failures
- Verify CSI driver advertises the expected capabilities
- Verify feature gates are enabled on kube-apiserver, kubelet, and external-attacher
- Check CSINode objects for missing topology/allocatable entries
Implementation History
- 2026-03-21: CSI spec PR #603 opened
- 2026-04-08: Discussed in SIG Storage meeting
- 2026-04-14: KEP drafted
Drawbacks
- Adds one new RPC and a request flag to the CSI spec, increasing spec surface area
- Requires coordination between kubelet and external-attacher during upgrade
- Adds one roundtrip to node registration (offset by lighter node-side call)
Alternatives
Alternative 1: Private CRD and Controller
A new CRD (CSINodeInfo) and controller to store and populate node info.
Why not: Replaces cloud API credentials with CR read permissions, which doesn’t fundamentally solve the security problem. Adds more roundtrips (cloud → controller → API server → node → kubelet → API server → scheduler). Harder to handle RESOURCE_EXHAUSTED. Kubernetes-specific, doesn’t help other COs.
Alternative 2: Instance Metadata Enhancement
Enhance cloud instance metadata services to provide all required information.
Why not: Metadata services typically provide only basic info (instance ID, zone). They don’t provide attachment limits, supported disk categories, or the list of CSI-managed volumes needed to infer non-CSI attachments. Requires coordination with every cloud provider. Not feasible for all environments.
Alternative 3: Node Label Patching (e.g., AWS metadata-labeler)
The AWS EBS CSI driver implements a metadata-labeler sidecar that runs on the controller, queries EC2 APIs for ENI and block device counts, and patches Node labels. The node-side driver reads these labels via the Kubernetes API.
The GCP PD CSI driver independently arrived at the same pattern:
a gce-pd-node-labeler
controller sidecar watches Node objects
and patches disk-type.gke.io/<type> labels indicating which disk types each machine family supports,
which the node-side driver reads back via the Kubernetes API to populate NodeGetInfoResponse’s accessible topology.
The Alibaba Cloud CSI driver did the same a third time:
a controller-side csi-metadata-labeler
queries ECS APIs for the disk categories and attach limit supported by each instance type,
then patches node.csi.alibabacloud.com/disktype.<type> labels and a max-disk annotation onto the Node;
the node plugin runs with --use-labeler=true to consume them.
Why not: Provider-specific, every driver needs its own solution — AWS, GCP, and Alibaba Cloud each independently reinvented the controller-side label-patching sidecar.
Mixes storage info into Node labels.
Not portable to other COs.
Cannot leverage VolumeAttachment objects to distinguish CSI-managed vs non-CSI volumes.
When the source metadata is unavailable and labels haven’t been patched yet, the node-side driver falls back to less accurate sources.
Alternative 4: CRD-based Topology Retrieval (e.g., vSphere CSINodeTopology)
The vSphere CSI driver uses a CSINodeTopology
CRD. The node creates a CR, the controller populates topology via vCenter API, and the node watches for completion before returning NodeGetInfoResponse.
Why not: Provider-specific. Node registration blocks until the controller updates the CR, and if the controller is slow or down, the node waits until timeout. Requires additional CRD and controller. Not portable to other COs.
Alternative 5: Hardcoded Instance-Type Tables in the Driver Binary
Because the node-side driver often lacks cloud credentials, several drivers ship static per-instance-type lookup tables compiled into the binary to answer max_volumes_per_node.
The AWS EBS CSI driver embeds a ~800-line table
that is regenerated from the EC2 DescribeInstanceTypes API at build time
.
The GCP PD CSI driver similarly hardcodes per-machine-family attach-limit tables
, but maintains them by hand.
Why not: Provider-specific, every driver needs its own table. The table is a frozen snapshot: it must be updated and the driver re-released for every new instance type or limit change, and a stale binary silently returns wrong limits. Notably, the underlying data is available from the cloud provider’s API (AWS regenerates its table from that very API) — the only reason it is baked into the binary is that the node lacks credentials to query it at runtime. A controller-side RPC, which runs where credentials already exist, can query the live API directly and eliminate the static table entirely. The Alibaba Cloud CSI driver already demonstrates this in practice: rather than hardcoding a table, its controller-side labeler queries ECS APIs at runtime for both the attach limit and the supported disk types of each instance type.
This proposal provides a standardized CSI spec approach that all drivers can adopt, avoiding provider-specific implementations.
Alternative 6: Separate NodeGetID RPC
Instead of a request flag on NodeGetInfo, add a dedicated NodeGetID RPC that returns only node_id.
The CO calls NodeGetID (not NodeGetInfo) when a driver advertises a GET_ID node capability, which serves the same handshake role as the flag.
This is a clean design — the response type structurally cannot carry topology or limits, so there is nothing for the CO to ignore. It was the original proposal for this KEP.
Why not:
Higher adoption cost: A new RPC must be implemented, tested, and versioned across the 100+ existing CSI drivers. The request flag requires only a capability plus one optional field, so drivers adopt it incrementally.
Two node-side RPCs with overlapping purpose:
NodeGetInfois not deprecated, so drivers and readers face two node identity RPCs whereNodeGetIDreturns a strict subset ofNodeGetInfo. This is confusing for new implementers, since CSI’s other capability-gated RPCs each add a distinct operation rather than a subset of an existing one.Duplicates shared node-produced fields:
node_idis produced by the node but consumed by the controller side, so it must live on a message the node returns. A future node-produced field needed by both flows (e.g. a node context relayed toControllerPublishVolume) has the same property. With the flag, such fields have one home onNodeGetInfoResponse; withNodeGetID, they must be duplicated on both messages and kept in sync.
The trade-off is that NodeGetInfoResponse now has two fields whose validity depends on the request flag.
Alternative 7: Combine Node and Controller Values
Instead of the controller being the sole authority for topology and limits when the flag is set,
let the node still report what it can (e.g. its zone from local metadata, or a coarse limit) and
have the values combined with the controller’s — for example min(node_count, controller_count) for the limit and a union of topology keys.
Why not:
No CO-agnostic combination rule:
minfor the limit and union for topology are one specific policy. Other drivers may wantmax, “controller wins unless zero”, or a per-key rule. The CO cannot know a driver’s intended semantic, so any rule the CO hard-codes is wrong for some driver.Correct combination belongs in the SP, which reintroduces the transport it was meant to avoid: doing it right means relaying the node’s partial values as input to
ControllerGetNodeInfoso the SP combines them. That growsDriverRegistrationsinto arbitrary node-side state and adds request fields and SP-side merge logic — significant complexity for a marginal gain.
The sole-authority design is simpler and has no rule-policy problem: the node reports only node_id, and the controller owns topology and limits.
Combination remains addable later if a concrete case shows the controller genuinely cannot determine a value the node can.
We may also consider the node context described in alternative 6.
Alternative 8: Reactive-Only Discovery
Skip upfront reporting entirely and let Kubernetes learn capacity from RESOURCE_EXHAUSTED failures.
Why not: Learning from failures is expensive — by the time an error occurs, Kubernetes has already created a volume (especially in WaitForFirstConsumer mode), scheduled a pod, and potentially started containers. Teardown involves expensive cloud operations. Additionally, this approach is one-sided: there is no mechanism to detect when capacity increases, only failures signal capacity decrease. Furthermore, most use cases do not have dynamic out-of-band attachment, so the effective volume limit is basically static. Designing a complex estimation algorithm (handling cold start, parallel attachment, multiple scheduler instances, etc.) to discover a static value is overkill.
However, reactive error handling remains necessary as a complement to proactive reporting. Out-of-band volume attachments (manually attached disks, network interfaces) can always occur, so the CO must handle violations after the fact. The existing RESOURCE_EXHAUSTED → re-query mechanism (KEP-4876) serves this purpose and is unchanged by this KEP.
Alternative 9: Static Node Context Object
Store controller-side information in a “node context” (similar to volume context or publish context) — call the controller once per node, persist the result, and have kubelet pass it to the node plugin on subsequent calls so the node can answer NodeGetInfo accurately.
Why not: This approach is fundamentally misaligned with the information flow. The CO must call the node first to obtain node_id, then call the controller with that ID. The controller’s output (topology, limits, attached volumes) is consumed by the CO itself for scheduling — there is no reason to route it back to the node plugin. The node plugin is not the consumer of this information; the scheduler and external-attacher are.
Additionally, even if we could pass controller context to the node, the data needed for accurate volume limit calculation (the list of attached volumes) is dynamic and constantly changing. A one-time-populate approach only works for static data, but the set-difference calculation for non-CSI volumes requires current state. This would require continuous polling and re-population, making it no simpler than the proposed design while adding an unnecessary extra hop.
Alternative 10: Controller Results in CSINode Status, Published by Kubelet
External-attacher could store results in CSINode status for kubelet to publish into spec.drivers and Node labels. This avoids controller Node patch permission, but requires an additional result representation, kubelet observation of CSINode, and another API publication step. Direct publication avoids that handoff; registration inputs and conditional writes coordinate the two writers.
Infrastructure Needed
- CSI spec update: PR #603
- CSI mock driver updated with the
controller_get_node_infoflag andControllerGetNodeInfoRPC for testing