KEP-6395: Dynamic Node-Local Ephemeral Volumes
KEP-6395: Dynamic Node-Local Ephemeral Volumes
- Release Signoff Checklist
- Summary
- Motivation
- Proposal
- Design Details
- Production Readiness Review Questionnaire
- Implementation History
- Drawbacks
- Alternatives
Release Signoff Checklist
Items marked with (R) are required prior to targeting to a milestone / release.
- (R) Enhancement issue in release milestone, which links to KEP dir in [kubernetes/enhancements] (not the initial KEP PR)
- (R) KEP approvers have approved the KEP status as
implementable - (R) Design details are appropriately documented
- (R) Test plan is in place, giving consideration to SIG Architecture and
SIG Testing input (including test refactors)
- e2e Tests for all features
- Tests run automatically on PRs
- (R) Graduation criteria is in place
- (R) Production readiness review completed
- (R) Production readiness review approved
- “Implementation History” section is up-to-date
- User-facing documentation has been created in [kubernetes/website], for publication to [kubernetes.io]
- Supporting documentation (e.g., additional design documents, links to mailing list discussions/SIG meetings, issues before this KEP) is included to the “Replaces” - “See Also” list
Summary
This KEP enables Kubernetes workloads to dynamically add and remove
node-local ephemeral volumes (configMap, secret, projected,
emptyDir, and OCI image) to and from pods via the /dynamic subresource
without requiring the entire pod to be destroyed and recreated.
For the Alpha release, volume mutations actuate upon:
- Container Restart: Modifying
volumeMountson an existing container via/dynamicexplicitly triggers Kubelet to recreate the container. - Container Addition: Mounting newly added or existing volumes into dynamically added containers (e.g., dynamic containers via KEP-5972: Dynamic Containers or ephemeral containers).
- Container Removal: Safely unmounting and tearing down volumes once all containers referencing them have terminated.
Live hot-plug into running containers without restart is an explicit non-goal
for Alpha and will be reevaluated in a future milestone. This enhancement requires no changes
to the Container Runtime Interface (CRI), does not interact with control-plane
storage controllers (AttachDetachController, PVController), and does not
require scheduler coordination.
Motivation
In Kubernetes today, a Pod’s storage configuration is strictly immutable once
scheduled. pod.Spec.Volumes cannot be altered, and container volumeMounts
cannot be updated. While this static model was sufficient for traditional
stateless microservices, it presents a severe bottleneck for modern workloads:
- Heterogeneous Task Workers & Batch Runners: Any workload architecture utilizing “worker” pods or containers that execute varying sequential tasks benefits from dynamic volume swapping. In batch processing frameworks, continuous delivery runners, and generic worker pools, pods remain running to amortize scheduling, admission, image pull, and sandbox runtime initialization costs. As different tasks arrive, each task demands its own configuration, credentials, or scratch datasets. Dynamic volumes allow worker pods to swap in the volumes required for each specific task without pod recreation.
- Sandboxed Workload Runners & Agentic Sandboxes: Modern sandboxed workload
runners host long-running containers (e.g., sandboxed runners) that execute
distinct jobs sequentially. Each job requires an isolated workspace volume
(e.g.,
emptyDirscratch or projected tokens). Because volumes are immutable, platforms are forced to choose between deleting and recreating pods for every job (incurring seconds to minutes of scheduling and runtime initialization overhead) or using insecure out-of-bandhostPathbind mounts that bypass Kubernetes security, quotas, and authorizers. - Pre-Warmed Sandbox Pods: Systems pre-warm pods on nodes with pre-pulled
container images and initialized runtimes. When a tenant workload arrives,
tenant-specific configuration (
configMap,secret,projected) must be injected into the pre-warmed pod without destroying the warm sandbox. - Dynamic Containers Synergy: Dynamic Containers introduces the
/dynamicsubresource to add and remove containers dynamically. However, Dynamic Containers explicitly constrains dynamic containers to only mount volumes already defined at pod creation. This proposal removes this limitation, allowing dynamic containers to bring their own volumes.
By allowing node-local volumes to be added and removed dynamically, workloads can mutate storage definitions natively through the Kubernetes API with sub-second turnaround times.
Goals
- Enable dynamically adding and removing node-local volumes (
configMap,secret,projected,emptyDir, and OCIimage) inpod.Spec.Volumeson running pods via the/dynamicsubresource. - Enable adding, removing, and updating
volumeMountsincontainer.VolumeMountsfor containers that are being restarted, added (ephemeral or dynamic containers), or removed. - Maintain complete backward compatibility for standard
/podsupdates, ensuring default immutability guarantees remain intact.
Non-Goals
- Live Hot-Plug into Running Containers (Alpha): Hot-mounting or hot-unmounting filesystems into currently running container processes without container restart is explicitly out of scope for Alpha. We will reevaluate this limitation for a future milestone or KEP.
- HostPath Volumes (Alpha): Dynamic addition of
hostPathvolumes is out of scope for Alpha due to security implications. We will reevaluate supportinghostPathvolumes for Beta. - PersistentVolumeClaims (PVCs): Dynamic addition and removal of PVCs, including local PVs and cloud block volumes, is out of scope and deferred to follow-up KEPs.
- Volume Resizing: Resizing existing volumes in-place is out of scope.
- Kube-Scheduler Changes: Inline node-local ephemeral volumes do not have persistent capacity tracking or node affinity; scheduling is unaffected.
Proposal
We propose extending the /dynamic subresource introduced in
KEP-5972: Dynamic Containers
to permit mutations to spec.volumes and container
spec.containers[*].volumeMounts for supported node-local volume types.
When a client submits an update to /dynamic:
- The API Server validates that newly added volumes are of supported types
(
configMap,secret,projected,emptyDir,image) and that any removed volumes are no longer referenced by active containers. - The Node Authorizer updates its graph, granting the assigned node authorization to read newly mounted Secrets or ConfigMaps.
- Kubelet admits the request through an Allocation step, recording the
admitted volume and container specifications into its internal state, which
is exposed via the
/allocatedsubresource. - Kubelet prepares and mounts the volume on the host filesystem.
- In the Actuation step, Kubelet mounts or unmounts volumes as needed, triggering a restart if volume mounts are modified on an existing container.
- When a volume is removed from
pod.Spec.Volumes, Kubelet tears down the volume mount on the host once all containers referencing it have stopped, after which the Node Authorizer graph removes authorization edges for those Secrets or ConfigMaps.
User Stories
Story 1: Autonomous Framework Governance & Workspace Swapping
An autonomous orchestrator or local runner maintains warm sandbox pods. As jobs
transition, the runner swaps out the tenant workspace. The runner submits a
request to /dynamic that removes the previous job’s emptyDir scratch
volume, adds a new emptyDir volume, and restarts the task container.
The container starts up immediately with the fresh volume, avoiding a full
pod recreation cycle.
Story 2: Pre-Warmed Sandbox Pods for Agentic Workloads
An agent platform pre-warms sandboxed runner pods on worker nodes. When an agent
requests a tool execution environment, the platform dynamically adds a
projected volume containing ephemeral credentials, tool configuration, and a
scratch disk, and launches a dynamic sidecar container referencing those mounts.
Execution begins in under 500ms.
Story 3: In-Place Sidecar & Workload Upgrades
A credential-rotation or logging sidecar needs to be upgraded or rotated. The
operator patches the pod via /dynamic to attach an updated ConfigMap volume
and restarts the sidecar container in-place, without disturbing the primary
application container running in the same pod.
Story 4: Heterogeneous Task Workers & Pipeline Runners
A long-running worker pod in a distributed data processing system or CI/CD
pipeline polls a queue for tasks. For Task 1, the orchestrator issues a
/dynamic update to mount a specific task ConfigMap and an emptyDir scratch
volume, triggering an explicit restart of the task worker container. When Task 1
completes, the orchestrator removes the scratch volume, mounts fresh task
credentials for Task 2, and continues execution within the same pod—eliminating
pod scheduling, CNI network setup, and container image pull latencies between
tasks.
Risks and Mitigations
Third-party controllers assuming static pod volumes
- Risk: Admission webhooks, GitOps agents, or third-party controllers may
assume that
.spec.volumesis immutable after pod creation. An unexpected update could cause panics or desynchronized state in external controllers. - Mitigation: Dynamic volume updates are strictly forbidden on the main
/podsendpoint and require explicit routing to the/dynamicsubresource (see KEP-5972: Dynamic Containers for further details on subresource access control and ecosystem migration). RBAC permissions on/dynamicare restricted to cluster administrators by default. The feature is gated behind an opt-in feature gate (DynamicNodeLocalEphemeralVolumes).
Resource exhaustion and eviction threshold interference
- Risk: Uncontrolled addition of disk-backed or memory-backed
emptyDirvolumes could exhaust node ephemeral storage or memory, triggering node-level evictions. - Mitigation: Memory-backed
emptyDirvolumes are accounted against container and pod memory limits. Disk-backedemptyDirvolumes are subject to project quotas where supported, and Kubelet’s existing ephemeral storage eviction manager monitors usage against node thresholds.
Dynamic Secret exfiltration and unmount (“Secret Scrubbing”)
- Risk: A compromised or malicious container temporarily mounts a sensitive
Secretvolume via/dynamic, copies the secret payload to an unencryptedemptyDirscratch volume or transmits it across the network, and immediately removes the Secret volume frompod.Spec.Volumesto conceal evidence of having accessed the secret from futurekubectl get podaudits. - Mitigation: Once container code executes, Kubernetes cannot prevent an
in-memory copy of secret data; therefore, protection relies on strict API
access boundaries and auditability. The Kubernetes API Audit Log captures
every
/dynamicsubresource request and patch chronologically with complete user identity, timestamp, and object payload, preserving an immutable record of every dynamically mounted and unmounted volume. Furthermore, RBAC permissions on/dynamicare restricted to cluster administrators and trusted controllers by default.
Desynchronized Secret lifetime during blocked host unmount
- Risk: A client removes a
Secretvolume frompod.Spec.Volumes, but unmount or teardown on the node is delayed or blocked (e.g., container restart is deferred or an open file descriptor holds the mount). During this interval, secret data remains physically resident on the node’s filesystem (/var/lib/kubelet/pods/<uid>/volumes/kubernetes.io~secret/<name>) even thoughspec.volumesno longer lists the secret. - Mitigation: Kubelet’s
/allocatedsubresource maintains the volume in the node’s admitted state until container termination and host directory unmount (TearDownAt) are fully verified. Observers inspecting/allocatedcan verify active node-level volume allocations. Furthermore, Kubelet secret volumes are mounted on RAM-backedtmpfsfilesystems; when Kubelet executesTearDownAt, unmounting thetmpfsimmediately zeroes and releases the in-memory secret payload.
Design Details
API Changes
This proposal introduces the following API changes:
- A new API field on
v1.Container:dynamicVolumePolicy(type*ContainerDynamicVolumePolicy), containingrestartPolicy(typeDynamicVolumeRestartPolicy). - Though not introduced by this proposal, we expand the
/dynamicsubresources introduced in KEP-5972: Dynamic Containers , expanding their schema allowances to permit node-local ephemeral volume mutations. - A new API field on
v1.PodStatus:volumeStatuses(type[]PodVolumeStatus), analogous tocontainerStatuses, to track whether volumes are allocated and prepared on the host node without coarse pod conditions (actuated container mounts continue to be reported underContainerStatuses[*].VolumeMounts).
Container Dynamic Volume Policy
To allow users to govern how a container behaves when its volumeMounts are
dynamically added or removed (at the same level as resizePolicy in KEP-1287),
a new dynamicVolumePolicy field is added to v1.Container:
spec:
containers:
- name: runner
image: registry.k8s.io/pause:3.10
dynamicVolumePolicy:
restartPolicy: RestartContainer
volumeMounts:
- name: scratch
mountPath: /mnt/scratch
In go:
type Container struct {
// ... existing fields ...
// DynamicVolumePolicy defines dynamic volume mutation behavior for this container.
// +featureGate=DynamicNodeLocalEphemeralVolumes
// +optional
DynamicVolumePolicy *ContainerDynamicVolumePolicy `json:"dynamicVolumePolicy,omitempty"`
}
// ContainerDynamicVolumePolicy defines dynamic volume mutation behavior for a container.
type ContainerDynamicVolumePolicy struct {
// RestartPolicy defines whether dynamic volume mount additions or removals
// require restarting the container. Supported values: RestartContainer, "".
// Defaults to "".
// +optional
RestartPolicy DynamicVolumeRestartPolicy `json:"restartPolicy,omitempty"`
}
// DynamicVolumeRestartPolicy defines the restart behavior applied to a container
// when its volume mounts are dynamically mutated.
// +enum
type DynamicVolumeRestartPolicy string
const (
// DynamicVolumeRestartContainer indicates that Kubelet must restart the container
// in-place to actuate volume mount additions or removals on a running container.
DynamicVolumeRestartContainer DynamicVolumeRestartPolicy = "RestartContainer"
)
The semantics of dynamicVolumePolicy.restartPolicy are defined as follows:
RestartContainer: SettingdynamicVolumePolicy.restartPolicy: RestartContainerpermits adding or removing volume mounts on an existing running container, and explicitly instructs Kubelet to restart (recreate) the container in-place to actuate the mount changes.- Default (""): When
dynamicVolumePolicyis omitted, the API server rejects requests to add or remove volume mounts while the container is running. Under an empty policy, volume mounts can only be added along with a container addition, and can only be removed along with a container removal. - Future (
NotRequired): In the future, we will consider supporting hot-plugging volumes to running containers without restart under an additional policy (e.g.NotRequired), but this is out of scope for the first alpha.
Integration with the /dynamic Subresource
KEP-5972: Dynamic Containers
introduces the /dynamic subresource. Dynamic volume mutations are executed
exclusively via PUT /api/v1/namespaces/{namespace}/pods/{name}/dynamic.
Dynamic volume mutations permit:
- Appending new volumes to
pod.Spec.Volumes. - Removing volumes from
pod.Spec.Volumesif they are no longer referenced by active containers. - Updating
volumeMountswithinpod.Spec.Containers[*]for containers undergoing restart or addition.
Any semantics regarding the /dynamic subresource that apply to dynamic containers
also apply to dynamic volumes.
Inspection via the /allocated Subresource
To provide observability into transactional state and decouple desired state
from admitted node state, this proposal integrates with the /allocated
subresource introduced by
KEP-5972: Dynamic Containers
.
The /allocated subresource is served directly by the Kubelet and represents
the current transactional allocation on the node. When a pod is updated via
/dynamic:
pod.Spec.Volumesreflects the desired volume configuration.- The
/allocatedpod endpoint reflects the admitted volume configuration that Kubelet has accepted and is actively preparing or maintaining on the node. - The already-existing
volumeMountsfield under container statuses (pod.Status.ContainerStatuses[*].VolumeMounts) reflects the actual actuated volume mounts in the running container.
Pod Volume Status API (pod.Status.VolumeStatuses)
To track the status of pod volumes across their allocation and host preparation
lifecycles, this proposal introduces volumeStatuses to v1.PodStatus.
Analogous to containerStatuses, this field provides granular, per-volume
observability into whether volumes are allocated and prepared on the host node.
Container mount actuation is not tracked in this field, but is surfaced via the
preexisting pod.Status.ContainerStatuses[*].VolumeMounts.
The allocated pod, including any dynamically added or removed volumes, is
exposed via the already-existing /allocated subresource on Kubelet.
Detailed failure reasons and diagnostics will be emitted via Events, which will be further designed in beta.
type PodStatus struct {
// ... existing fields ...
// VolumeStatuses contains the status of volumes in the pod, tracking whether
// volumes are allocated and prepared on the host node.
// +listType=map
// +listMapKey=name
// +optional
// +featureGate=DynamicNodeLocalEphemeralVolumes
VolumeStatuses []PodVolumeStatus `json:"volumeStatuses,omitempty"`
}
// PodVolumeStatus represents the status of an individual volume in a pod.
type PodVolumeStatus struct {
// Name is the name of the volume matching pod.Spec.Volumes.
Name string `json:"name"`
// HostPhase represents the host-level preparation phase of the volume:
// Pending, HostMounted, ErrorMounting, Unmounting.
// +optional
HostPhase PodVolumeHostPhase `json:"hostPhase,omitempty"`
// ErrorMessage represents the error encountered during the host preparation
// phase of the volume. It is only populated when an error occurs during SetUpAt.
// +optional
ErrorMessage string `json:"errorMessage,omitempty"`
}
// PodVolumeHostPhase represents the host preparation phase of a volume mounted to a pod.
type PodVolumeHostPhase string
const (
// PodVolumeHostPending indicates that the volume has been added to pod.Spec.Volumes,
// and is waiting node allocation or host preparation/mounting by VolumeManager.
PodVolumeHostPending PodVolumeHostPhase = "Pending"
// PodVolumeHostMounted indicates that VolumeManager has successfully prepared and
// mounted the volume on the node host (marked ready in ActualStateOfWorld).
PodVolumeHostMounted PodVolumeHostPhase = "HostMounted"
// PodVolumeHostErrorMounting indicates that Kubelet admitted the volume into allocatedPod,
// but VolumeManager encountered an error during SetUpAt while preparing or mounting the
// volume on the host.
PodVolumeHostErrorMounting PodVolumeHostPhase = "ErrorMounting"
// PodVolumeHostUnmounting indicates that volume removal has been admitted and allocated
// by Kubelet, and Kubelet is actively unmounting container bind mounts and executing host TearDownAt.
PodVolumeHostUnmounting PodVolumeHostPhase = "Unmounting"
)
Lifecycle Phases
Pending: A dynamic volume mutation was submitted inpod.Spec.Volumes, but Kubelet has not yet admitted it intoallocatedPod(for example, due to node ephemeral storage quota limits or a rapid re-add collision). If admission fails, the volume remains inPending, and Kubelet emits a corresponding event (e.g.FailedAllocation).HostMounted: Kubelet has admitted the volume intoallocatedPod,VolumeManagerhas successfully executedSetUpAton the node host, and the volume path (/var/lib/kubelet/pods/<uid>/volumes/...) is recorded ready inActualStateOfWorld(ASW).ErrorMounting: Kubelet admitted the volume intoallocatedPod, butVolumeManagerencountered an error duringSetUpAtwhile preparing or mounting the volume on the host (for example, failure to resolve a referenced Secret or ConfigMap, an image pull failure, or a filesystem mount syscall failure). This ensures failures during the host mounting step are explicitly surfaced rather than leaving the volume inPending. Detailed error context is recorded in Kubelet Events and in theErrorMessagefield.Unmounting: The volume removal frompod.Spec.Volumeshas been successfully admitted and committed toallocatedPod.Unmountingis set after the removal has been admitted and allocated by Kubelet. If a volume removal is part of a transactional dynamic update that is rejected or deferred during Kubelet allocation, the volume removal is not committed toallocatedPod, and itshostPhaseremainsHostMounted. Once allocation succeeds, Kubelet retains the volume entry inUnmountingwhile target containers unmount the volume andVolumeManagerexecutes hostTearDownAt. Once host unmount completes and the volume is cleared from ASW, the entry is pruned fromVolumeStatuses.
Sample pod.Status Lifecycle States
State 1: Pending Volume and Container
When a dynamic mutation adding a new volume and a container mounting it is submitted
via /dynamic, but cannot yet be admitted by Kubelet:
status:
volumeStatuses:
- name: dynamic-scratch
hostPhase: Pending
containerStatuses:
- name: worker
state:
waiting:
reason: Unallocated
message: "Container addition pending node allocation"
State 2: Host Mounted, Container Restart Pending
When a dynamic mutation adding a new volume and mounting it into an existing container is submitted, Kubelet admits the volume into allocatedPod. VolumeManager mounts the host
directory in ASW (hostPhase: HostMounted). The following yaml shows the pod status after this step, but before the target container has restarted to pick up the new mount:
status:
volumeStatuses:
- name: dynamic-scratch
hostPhase: HostMounted
containerStatuses:
- name: runner
ready: true
state:
running:
startedAt: "2026-09-25T14:00:00Z"
# volumeMounts still reflects the pre-restart container mounts
volumeMounts: []
State 3: Fully Actuated
Kubelet recreates container runner with the volume mount. The container runtime actuates the bind mount into the container namespace, and pod.Status.ContainerStatuses[*].VolumeMounts is updated:
status:
volumeStatuses:
- name: dynamic-scratch
hostPhase: HostMounted
containerStatuses:
- name: runner
ready: true
state:
running:
startedAt: "2026-09-25T14:05:00Z"
volumeMounts:
- name: dynamic-scratch
mountPath: /mnt/scratch
readOnly: false
State 4: Host Mount Failure (ErrorMounting)
If Kubelet admits the volume into allocatedPod, but VolumeManager encounters an error during SetUpAt on the node:
status:
volumeStatuses:
- name: dynamic-scratch
hostPhase: ErrorMounting
containerStatuses:
- name: runner
ready: true
state:
running:
startedAt: "2026-09-25T14:00:00Z"
volumeMounts: []
API Server Validation Rules
When a client issues a PUT to /dynamic, the API server validates the
request before persisting the pod specification to etcd:
- Dynamic Volume Policy Validation:
- Adding or removing volume mounts on an already-running container requires
dynamicVolumePolicy.restartPolicy: RestartContaineron that container. If a request attempts to add, remove, or mutate mounts on a running container with an empty policy, the update is rejected with a validation error. - Adding mounts with an empty restart policy is permitted only when simultaneously adding a new container (such as a dynamic or ephemeral container).
- Adding or removing volume mounts on an already-running container requires
- Exclusion of Init Containers (Alpha): Dynamic volume mutations targeting
pod.Spec.InitContainers[*]are strictly rejected in Alpha. Any update attempting to add, remove, or mutatevolumeMountson init containers returns a validation error. Supporting init containers (including restartable sidecar init containers) will be reevaluated for Beta. - Supported Volume Sources: Only node-local ephemeral volume types are
permitted:
v1.VolumeSource.ConfigMapv1.VolumeSource.Secretv1.VolumeSource.Projectedv1.VolumeSource.EmptyDirv1.VolumeSource.Image
- Unsupported Types: Updates attempting to add
PersistentVolumeClaim,HostPath,CSI, or cloud volume sources are rejected with a validation error. - Volume Name Uniqueness and Collision Prevention:
- API Server Validation: Volume names in
spec.volumesmust remain unique. In addition, the API server rejects adding a new volume if its name is currently present inpod.Status.VolumeStatusesor still actively reported in any container’spod.Status.ContainerStatuses[*].VolumeMounts. This prevents race conditions where a client removes a volume and rapidly re-adds one with the same name before container unmounting and host teardown complete, exactly mirroring Dynamic Containers.
- API Server Validation: Volume names in
- Dangling Mount Protection: A volume cannot be deleted from
spec.volumesif any container inspec.containersstill mounts it, unless that container is also being removed or updated to remove the mount in the same transaction. - Pod Phase: Dynamic volume mutations are only allowed while the pod is in
the
Runningphase and has completed initialization. OnceDeletionTimestampis set, mutations are rejected.
When pod volumes are updated via /dynamic, the Node Authorizer’s existing
informer automatically updates its graph to grant the node access to newly
mounted Secrets or ConfigMaps. Any transient propagation delay is safely handled
by Kubelet’s volume mounting retry loop.
Two-Stage Kubelet Lifecycle: Allocation and Actuation
Mutations submitted through the /dynamic subresource follow a two-stage
Kubelet lifecycle: Allocation and Actuation. While the /dynamic
subresource itself comes from
KEP-5972: Dynamic Containers
,
splitting node admission from container runtime configuration adopts the
two-stage allocation and actuation model established in
KEP-1287: In-Place Pod Vertical Scaling
. Once the API server
validates the mutation and persists it to etcd, both stages are executed
locally on the node by Kubelet.
+-----------------------------------------------------------------------------+
| Stage 1: Allocation (Kubelet Admission & Async Host Preparation) |
| |
| 1. Kubelet admits volume mutation against node-local resources and quotas. |
| 2. Kubelet checkpoints admitted state to allocatedPod. |
| 3. Kubelet reflects admitted volumes in the /allocated pod representation. |
| 4. VolumeManager DSWP discovers volume in allocatedPod and mounts on host |
| asynchronously via SetUpAt in the background. |
+-----------------------------------------------------------------------------+
│
▼
+-----------------------------------------------------------------------------+
| Stage 2: Actuation (Container-Scoped Runtime Configuration) |
| |
| 1. Initial SyncPod (Pod Startup): Blocks on WaitForAttachAndMount before |
| starting any initial containers (preserving standard Kubernetes behavior)|
| 2. Subsequent SyncPods (Dynamic Mutations): Kubelet checks if volumes |
| required by target containers are ready in ASW. |
| 3. If a dynamic volume is still mounting, only the specific target container|
| restart/addition is deferred; unaffected containers continue running. |
| 4. Once ready, Kubelet recreates the target container (or starts the dynamic|
| container) with bind mounts pointing to the prepared host paths. |
| 5. Kubelet updates pod.Status.VolumeStatuses (hostPhase) and CRI mounts. |
+-----------------------------------------------------------------------------+
Stage 1: Admission and Allocation (Node-Local)
When a pod update arrives on the node via watch from the API server:
- Kubelet Admission: Kubelet validates that referenced Secrets and
ConfigMaps are resolvable and that local ephemeral storage or memory quota is
sufficient for requested
emptyDirvolumes. - State Checkpointing: Kubelet records the admitted volume and container
specifications into its local state checkpoint (
allocatedPod). - Subresource Exposure: Kubelet reflects the admitted configuration in the
/allocatedsubresource, confirming to external observers that allocation succeeded. - VolumeManager Asynchronous Host Preparation:
Kubelet’s
VolumeManageris only ever aware of theallocatedPod. TheDesiredStateOfWorldPopulatorreads volumes exclusively fromallocatedPodrather than unadmitted pod specs. This guarantees thatVolumeManagernever prepares or mounts host directories for mutations that failed local admission.VolumeManagergenerates the desired state of world (DSW) entry and invokesSetUpAton the appropriate volume plugin asynchronously in the background.
Stage 2: Actuation (Container Runtime)
Once host allocation and directory preparation are underway, volumes are
actuated into target containers during Kubelet’s SyncPod execution:
- Initial Pod Startup vs. Subsequent Dynamic SyncPods:
- Initial Pod Startup: The very first
SyncPodinvocation for a new pod preserves standard Kubernetes behavior: Kubelet blocks onWaitForAttachAndMountto ensure all creation-time volumes are fully mounted on the host before any containers or init containers start. - Subsequent Dynamic SyncPods: For already-running pods processing
dynamic mutations, Kubelet avoids globally blocking or stalling the pod
worker. Instead, during container reconciliation (
computePodActions), Kubelet inspectsActualStateOfWorld(ASW) for the specific volumes required by each container.
- Initial Pod Startup: The very first
- Container-Scoped Deferral & Volume Readiness Trigger:
- Determining Readiness via ASW: Kubelet determines volume readiness by
checking VolumeManager’s
ActualStateOfWorld(ASW) cache. - Container-Scoped Deferral: If a newly added volume required by a container has not yet been marked mounted in ASW, Kubelet defers actuation only for that specific container. Unaffected containers continue running uninterrupted.
- Event-Driven Trigger: When VolumeManager finishes mounting a volume and marks it in ASW, it triggers a pod sync to immediately actuate the deferred container.
- Determining Readiness via ASW: Kubelet determines volume readiness by
checking VolumeManager’s
- Explicit Container Re-creation Decisions:
When the required volumes are verified ready in ASW,
computePodActionsdetects that an existing container’svolumeMountsinallocatedPoddiffer from the running container configuration and schedules an explicit container restart (stop and recreate). For newly introduced dynamic or ephemeral containers, Kubelet schedules container creation. - CRI Invocation:
Kubelet stops the existing container if restarting, constructs the CRI
container specification with bind mounts pointing to the prepared host paths,
and invokes CRI
CreateContainerfollowed byStartContainer. - Status and Actuation Reflection:
Kubelet updates
pod.Status.VolumeStatusesto reflect host-level volume readiness (hostPhase: HostMounted, orErrorMountingon failure). Actuation into containers is reflected exclusively underpod.Status.ContainerStatuses[*].VolumeMountsas CRI confirms the active bind mounts. User observability into allocation details and failure causes is provided via Kubelet Events and the/allocatedsubresource rather than bloatingPodStatus.
Volume Teardown and Removal Lifecycle
Volume removal follows the same two-stage allocation and actuation pattern, ensuring that container unmounting precedes host directory teardown:
- API Validation: The API server strictly rejects removing a volume from
spec.volumesif any container inspec.containersstill mounts it, unless that container mount is also removed in the same/dynamicupdate. - Allocation Stage (Atomic Admission):
When a
/dynamicupdate removes a volume along with its referencing container (or removes the mount from an existing container), Kubelet admits the mutation and commits the changes toallocatedPodatomically. - Actuation Stage (Container Unmount):
During actuation, Kubelet’s
SyncPodacts on the container first, stopping and recreating it with the updated container specification (excluding the removed volume mount). The container runtime releases the mount from the container’s mount namespace. - Teardown & Host Unmount via
PodStateProvider:VolumeManager’sDesiredStateOfWorldPopulator(DSWP) observes that the volume has been removed fromallocatedPod. To prevent tearing down host directories while containers are still running, DSWP queries Kubelet via an extension to thePodStateProviderinterface:Kubelet implementstype PodStateProvider interface { // IsVolumeInUseByPod returns true if any running container in the pod // currently mounts the specified volume. IsVolumeInUseByPod(podUID types.UID, volumeName string) bool }IsVolumeInUseByPodby checking its internal, in-memory rutime cache of running container statuses. Once CRI confirms the old container has stopped, the runtime cache clears the mount reference. Only then does DSWP remove the volume fromDesiredStateOfWorld(DSW). The reconciler then unmounts and remove the host directory. - Status Pruning:
When a volume removal is submitted via
/dynamic, its entry inpod.Status.VolumeStatusestransitions tohostPhase: Unmountingonly after the removal has been successfully allocated by Kubelet. If the removal is part of a transactional dynamic update that is rejected or deferred during Kubelet allocation (for example, if submitted alongside an infeasible container resource resize), the volume removal is not committed toallocatedPod, and itshostPhaseremainsHostMounted. Once allocated, Kubelet retainsUnmountingwhile containers are restarted or stopped without the volume mount (updatingpod.Status.ContainerStatuses[*].VolumeMounts). OncePodStateProvider.IsVolumeInUseByPodreturns false and VolumeManager completes hostTearDownAt(clearing ASW), Kubelet prunes the volume entry completely frompod.Status.VolumeStatuses.
Limitations
- Actuation occurs only upon container restart, addition (ephemeral or dynamic containers), or removal.
- Live hot-plug into running containers without restart (
NotRequired) is deferred to a future milestone or enhancement. - Dynamic volume mutations for
initContainers(including restartable sidecar init containers) are not supported in Alpha and will be reevaluated for Beta. - Dynamic addition of
hostPathvolumes is not supported in Alpha due to security considerations and will be reevaluated for Beta.
Test Plan
[x] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.
Prerequisite testing updates
None.
Unit tests
- API validation:
- Adding supported node-local volume types (
configMap,secret,emptyDir,projected,image) via/dynamic. - Rejection of unsupported types (
PVC,hostPath,CSI). - Rejection of duplicate volume names and dangling container mounts.
- Adding supported node-local volume types (
- Node Authorizer:
- Graph edge creation for dynamically mounted Secrets and ConfigMaps on running pods.
- Kubelet VolumeManager populator:
- Discovery of newly added volumes from
allocatedPodon running pods. - Teardown of removed volumes while containers remain running.
- Discovery of newly added volumes from
- Volume plugins (
configmap,secret,emptydir,projected):- Setup and teardown lifecycle on dynamic add/remove.
Integration tests
- API Server
/dynamicsubresource endpoint:- RBAC enforcement (cluster-admin required).
- Updates to
spec.volumesandvolumeMountssucceed.
- Node Authorizer integration:
- Kubelet can retrieve Secrets mounted dynamically into a running pod.
e2e tests
- Dynamic
emptyDiron container restart:- Start a pod, dynamically add an
emptyDirvolume, restart the container in-place, and verify the container writes data to the new mount.
- Start a pod, dynamically add an
- Dynamic
ConfigMapandSecretaddition:- Add a
ConfigMapandSecretvolume dynamically, restart the container, and verify keys are mapped to files inside the container.
- Add a
- Dynamic volume removal:
- Remove a volume from the pod spec, verify the host directory is unmounted and cleaned up by Kubelet.
- Synergy with
KEP-5972: Dynamic Containers
:
- Add a dynamic container and a new volume simultaneously via
/dynamic; verify container starts successfully with the volume mounted.
- Add a dynamic container and a new volume simultaneously via
Graduation Criteria
Alpha
- Feature gate
DynamicNodeLocalEphemeralVolumesimplemented (default disabled). dynamicVolumePolicyfield added tov1.Container(with nestedrestartPolicy).- Support for node-local volumes (
configMap,secret,projected,emptyDir,image) via/dynamic. - Allocation and actuation on container restart, container addition, and container removal.
- Unit, integration, and e2e tests passing.
Beta
- Consider enabling live hot-plug into running containers, relaxing validation on
container.dynamicVolumePolicy.restartPolicyto permitNotRequiredfor live mutations. This may be added in alpha2 or a separate KEP. - Reevaluate supporting
hostPathvolumes with appropriate security safeguards. - Reevaluate dynamic volume mutations for
initContainers(including restartable sidecar init containers). - Explore the feasibility of allowing Kubelet to retry admitting a rapidly re-added volume after the previous volume has completed host unmount and cleared ASW.
- Metrics for dynamic volume mount and unmount latency and error rates.
- Evaluate integration with
pod.status.volumeHealth(KEP-1432: Volume Health Monitor) asCSIVolumeHealthstabilizes, particularly for reporting underlying volume faults alongside dynamic volume lifecycle states.
GA
- Allowing time for feedback, with at least 2 release cycles in Beta / enabled by default.
- Further GA criteria to be added in the beta update.
Upgrade / Downgrade Strategy
- Upgrade:
- Enabling the feature gate introduces the
/dynamicvolume capability. Existing pods are unaffected.
- Enabling the feature gate introduces the
- Downgrade:
- Disabling the feature gate restores strict immutability. Pods that
previously added volumes continue running with their current mounts, but
further
/dynamicupdates to volumes are rejected.
- Disabling the feature gate restores strict immutability. Pods that
previously added volumes continue running with their current mounts, but
further
Version Skew Strategy
Version skew will be handled by NodeDeclaredFeatures. If a client issues an
update request to a Node that does not support it, the API server will reject the request.
Production Readiness Review Questionnaire
Feature Enablement and Rollback
How can this feature be enabled / disabled in a live cluster?
- Feature Gate:
DynamicNodeLocalEphemeralVolumes - Components:
kube-apiserver,kubelet
Does enabling the feature change any default behavior?
Standard updates to /pods remain immutable, but third-party controllers
and ecosystem tools may assume that a pod’s volumes field is stable and
immutable after creation. Enabling this feature changes that assumption, as
spec.volumes can now be mutated dynamically via the /dynamic subresource.
The /dynamic subresource was intended to capture the intent that “all fields may be made mutable under this subresource”,
so using it for this purpose should not be surprising to users.
Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?
Yes. Disabling the feature gate prevents any further dynamic volume mutations via the API.
What happens if we reenable the feature if it was previously rolled back?
The API server resumes accepting dynamic volume updates on /dynamic.
Are there any tests for feature enablement/disablement?
Yes. Unit and integration tests will verify that disabling the feature gate rejects dynamic volume updates with a feature disabled validation error.
Rollout, Upgrade and Rollback Planning
How can a rollout or rollback fail? Can it impact already running workloads?
What specific metrics should inform a rollback?
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?
Monitoring Requirements
How can an operator determine if the feature is in use by workloads?
How can someone using this feature know that it is working for their instance?
- Events
- Event Reason:
- API .status
- Condition name:
- Other field:
- Other (treat as last resort)
- Details:
What are the reasonable SLOs (Service Level Objectives) for the enhancement?
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
- Metrics
- Metric name:
- [Optional] Aggregation method:
- Components exposing the metric:
- Other (treat as last resort)
- Details:
Are there any missing metrics that would be useful to have to improve observability of this feature?
Dependencies
Does this feature depend on any specific services running in the cluster?
Scalability
Will enabling / using this feature result in any new API calls?
Yes, updates to /dynamic generate write requests and pod update watch events.
Will enabling / using this feature result in introducing new API types?
No, there are new API types.
Will enabling / using this feature result in any new calls to the cloud provider?
No. Node-local volumes have no cloud provider interactions.
Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?
No. Dynamic volume mutation is a new operation not covered by existing SLIs/SLOs.
Will enabling / using this feature result in increasing size or count of the existing API objects?
Modest increase in the size of v1.Pod objects when additional volume entries
are declared or when the volume restart policy is set.
Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?
Negligible CPU and memory impact on Kubelet. Disk usage corresponds to files
created in emptyDir.
Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?
No. Dynamic volumes do not affect this kind of node resources.
Troubleshooting
How does this feature react if the API server and/or etcd is unavailable?
What are other known failure modes?
What steps should be taken if SLOs are not being met to determine the problem?
Implementation History
- 2026-09-21: KEP created for Alpha.
Drawbacks
- Allowing
.spec.volumesto mutate departs from Kubernetes’’’ historical assumption of static pod specs, requiring ecosystem controllers to adjust their reconciliation models.
Alternatives
Alternative 1: Out-of-band HostPath Bind Mounts
Workloads mount a shared host directory and create subdirectories manually.
- Why rejected: Bypasses Kubernetes RBAC, volume authorization, quotas, and security boundaries.
Alternative 2: Full Pod Recreation
Destroy and recreate the pod whenever storage requirements change.
- Why rejected: Unacceptable latency (seconds to minutes) for high-frequency agentic and sandboxed workloads.
Alternative 3: Dedicated Dynamic Volume Container Type
Introduce a new volume type (e.g. dynamicVolumePool) that contains a mutable
list of nested volumes.
- Why rejected: Adds significant API surface complexity. Mutating
spec.volumesdirectly via/dynamicprovides a cleaner, uniform experience that naturally mirrors container dynamism in KEP-5972: Dynamic Containers .
Alternative 4: Live Hot-Plug into Running Containers for Alpha
Support live in-place hot-mounting into running containers without restart in Alpha via mount propagation.
- Why rejected for Alpha: While technically feasible, it adds operational complexity (requiring parent mount setup and specific directory conventions). Deferring to Beta allows the core API and VolumeManager lifecycle to stabilize first.