KEP-6395: Dynamic Node-Local Ephemeral Volumes

Implementation History
ALPHA Implementable
Created 2026-09-21
Latest v1.38
Milestones
Alpha v1.38
Beta TBD
Stable TBD
Ownership
Owning SIG
SIG Storage
Participating SIGs
Primary Authors

KEP-6395: Dynamic Node-Local Ephemeral Volumes

Release Signoff Checklist

Items marked with (R) are required prior to targeting to a milestone / release.

  • (R) Enhancement issue in release milestone, which links to KEP dir in [kubernetes/enhancements] (not the initial KEP PR)
  • (R) KEP approvers have approved the KEP status as implementable
  • (R) Design details are appropriately documented
  • (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
    • e2e Tests for all features
    • Tests run automatically on PRs
  • (R) Graduation criteria is in place
  • (R) Production readiness review completed
  • (R) Production readiness review approved
  • “Implementation History” section is up-to-date
  • User-facing documentation has been created in [kubernetes/website], for publication to [kubernetes.io]
  • Supporting documentation (e.g., additional design documents, links to mailing list discussions/SIG meetings, issues before this KEP) is included to the “Replaces” - “See Also” list

Summary

This KEP enables Kubernetes workloads to dynamically add and remove node-local ephemeral volumes (configMap, secret, projected, emptyDir, and OCI image) to and from pods via the /dynamic subresource without requiring the entire pod to be destroyed and recreated.

For the Alpha release, volume mutations actuate upon:

  1. Container Restart: Modifying volumeMounts on an existing container via /dynamic explicitly triggers Kubelet to recreate the container.
  2. Container Addition: Mounting newly added or existing volumes into dynamically added containers (e.g., dynamic containers via KEP-5972: Dynamic Containers or ephemeral containers).
  3. Container Removal: Safely unmounting and tearing down volumes once all containers referencing them have terminated.

Live hot-plug into running containers without restart is an explicit non-goal for Alpha and will be reevaluated in a future milestone. This enhancement requires no changes to the Container Runtime Interface (CRI), does not interact with control-plane storage controllers (AttachDetachController, PVController), and does not require scheduler coordination.

Motivation

In Kubernetes today, a Pod’s storage configuration is strictly immutable once scheduled. pod.Spec.Volumes cannot be altered, and container volumeMounts cannot be updated. While this static model was sufficient for traditional stateless microservices, it presents a severe bottleneck for modern workloads:

  1. Heterogeneous Task Workers & Batch Runners: Any workload architecture utilizing “worker” pods or containers that execute varying sequential tasks benefits from dynamic volume swapping. In batch processing frameworks, continuous delivery runners, and generic worker pools, pods remain running to amortize scheduling, admission, image pull, and sandbox runtime initialization costs. As different tasks arrive, each task demands its own configuration, credentials, or scratch datasets. Dynamic volumes allow worker pods to swap in the volumes required for each specific task without pod recreation.
  2. Sandboxed Workload Runners & Agentic Sandboxes: Modern sandboxed workload runners host long-running containers (e.g., sandboxed runners) that execute distinct jobs sequentially. Each job requires an isolated workspace volume (e.g., emptyDir scratch or projected tokens). Because volumes are immutable, platforms are forced to choose between deleting and recreating pods for every job (incurring seconds to minutes of scheduling and runtime initialization overhead) or using insecure out-of-band hostPath bind mounts that bypass Kubernetes security, quotas, and authorizers.
  3. Pre-Warmed Sandbox Pods: Systems pre-warm pods on nodes with pre-pulled container images and initialized runtimes. When a tenant workload arrives, tenant-specific configuration (configMap, secret, projected) must be injected into the pre-warmed pod without destroying the warm sandbox.
  4. Dynamic Containers Synergy: Dynamic Containers introduces the /dynamic subresource to add and remove containers dynamically. However, Dynamic Containers explicitly constrains dynamic containers to only mount volumes already defined at pod creation. This proposal removes this limitation, allowing dynamic containers to bring their own volumes.

By allowing node-local volumes to be added and removed dynamically, workloads can mutate storage definitions natively through the Kubernetes API with sub-second turnaround times.

Goals

  • Enable dynamically adding and removing node-local volumes (configMap, secret, projected, emptyDir, and OCI image) in pod.Spec.Volumes on running pods via the /dynamic subresource.
  • Enable adding, removing, and updating volumeMounts in container.VolumeMounts for containers that are being restarted, added (ephemeral or dynamic containers), or removed.
  • Maintain complete backward compatibility for standard /pods updates, ensuring default immutability guarantees remain intact.

Non-Goals

  • Live Hot-Plug into Running Containers (Alpha): Hot-mounting or hot-unmounting filesystems into currently running container processes without container restart is explicitly out of scope for Alpha. We will reevaluate this limitation for a future milestone or KEP.
  • HostPath Volumes (Alpha): Dynamic addition of hostPath volumes is out of scope for Alpha due to security implications. We will reevaluate supporting hostPath volumes for Beta.
  • PersistentVolumeClaims (PVCs): Dynamic addition and removal of PVCs, including local PVs and cloud block volumes, is out of scope and deferred to follow-up KEPs.
  • Volume Resizing: Resizing existing volumes in-place is out of scope.
  • Kube-Scheduler Changes: Inline node-local ephemeral volumes do not have persistent capacity tracking or node affinity; scheduling is unaffected.

Proposal

We propose extending the /dynamic subresource introduced in KEP-5972: Dynamic Containers to permit mutations to spec.volumes and container spec.containers[*].volumeMounts for supported node-local volume types.

When a client submits an update to /dynamic:

  1. The API Server validates that newly added volumes are of supported types (configMap, secret, projected, emptyDir, image) and that any removed volumes are no longer referenced by active containers.
  2. The Node Authorizer updates its graph, granting the assigned node authorization to read newly mounted Secrets or ConfigMaps.
  3. Kubelet admits the request through an Allocation step, recording the admitted volume and container specifications into its internal state, which is exposed via the /allocated subresource.
  4. Kubelet prepares and mounts the volume on the host filesystem.
  5. In the Actuation step, Kubelet mounts or unmounts volumes as needed, triggering a restart if volume mounts are modified on an existing container.
  6. When a volume is removed from pod.Spec.Volumes, Kubelet tears down the volume mount on the host once all containers referencing it have stopped, after which the Node Authorizer graph removes authorization edges for those Secrets or ConfigMaps.

User Stories

Story 1: Autonomous Framework Governance & Workspace Swapping

An autonomous orchestrator or local runner maintains warm sandbox pods. As jobs transition, the runner swaps out the tenant workspace. The runner submits a request to /dynamic that removes the previous job’s emptyDir scratch volume, adds a new emptyDir volume, and restarts the task container. The container starts up immediately with the fresh volume, avoiding a full pod recreation cycle.

Story 2: Pre-Warmed Sandbox Pods for Agentic Workloads

An agent platform pre-warms sandboxed runner pods on worker nodes. When an agent requests a tool execution environment, the platform dynamically adds a projected volume containing ephemeral credentials, tool configuration, and a scratch disk, and launches a dynamic sidecar container referencing those mounts. Execution begins in under 500ms.

Story 3: In-Place Sidecar & Workload Upgrades

A credential-rotation or logging sidecar needs to be upgraded or rotated. The operator patches the pod via /dynamic to attach an updated ConfigMap volume and restarts the sidecar container in-place, without disturbing the primary application container running in the same pod.

Story 4: Heterogeneous Task Workers & Pipeline Runners

A long-running worker pod in a distributed data processing system or CI/CD pipeline polls a queue for tasks. For Task 1, the orchestrator issues a /dynamic update to mount a specific task ConfigMap and an emptyDir scratch volume, triggering an explicit restart of the task worker container. When Task 1 completes, the orchestrator removes the scratch volume, mounts fresh task credentials for Task 2, and continues execution within the same pod—eliminating pod scheduling, CNI network setup, and container image pull latencies between tasks.

Risks and Mitigations

Third-party controllers assuming static pod volumes

  • Risk: Admission webhooks, GitOps agents, or third-party controllers may assume that .spec.volumes is immutable after pod creation. An unexpected update could cause panics or desynchronized state in external controllers.
  • Mitigation: Dynamic volume updates are strictly forbidden on the main /pods endpoint and require explicit routing to the /dynamic subresource (see KEP-5972: Dynamic Containers for further details on subresource access control and ecosystem migration). RBAC permissions on /dynamic are restricted to cluster administrators by default. The feature is gated behind an opt-in feature gate (DynamicNodeLocalEphemeralVolumes).

Resource exhaustion and eviction threshold interference

  • Risk: Uncontrolled addition of disk-backed or memory-backed emptyDir volumes could exhaust node ephemeral storage or memory, triggering node-level evictions.
  • Mitigation: Memory-backed emptyDir volumes are accounted against container and pod memory limits. Disk-backed emptyDir volumes are subject to project quotas where supported, and Kubelet’s existing ephemeral storage eviction manager monitors usage against node thresholds.

Dynamic Secret exfiltration and unmount (“Secret Scrubbing”)

  • Risk: A compromised or malicious container temporarily mounts a sensitive Secret volume via /dynamic, copies the secret payload to an unencrypted emptyDir scratch volume or transmits it across the network, and immediately removes the Secret volume from pod.Spec.Volumes to conceal evidence of having accessed the secret from future kubectl get pod audits.
  • Mitigation: Once container code executes, Kubernetes cannot prevent an in-memory copy of secret data; therefore, protection relies on strict API access boundaries and auditability. The Kubernetes API Audit Log captures every /dynamic subresource request and patch chronologically with complete user identity, timestamp, and object payload, preserving an immutable record of every dynamically mounted and unmounted volume. Furthermore, RBAC permissions on /dynamic are restricted to cluster administrators and trusted controllers by default.

Desynchronized Secret lifetime during blocked host unmount

  • Risk: A client removes a Secret volume from pod.Spec.Volumes, but unmount or teardown on the node is delayed or blocked (e.g., container restart is deferred or an open file descriptor holds the mount). During this interval, secret data remains physically resident on the node’s filesystem (/var/lib/kubelet/pods/<uid>/volumes/kubernetes.io~secret/<name>) even though spec.volumes no longer lists the secret.
  • Mitigation: Kubelet’s /allocated subresource maintains the volume in the node’s admitted state until container termination and host directory unmount (TearDownAt) are fully verified. Observers inspecting /allocated can verify active node-level volume allocations. Furthermore, Kubelet secret volumes are mounted on RAM-backed tmpfs filesystems; when Kubelet executes TearDownAt, unmounting the tmpfs immediately zeroes and releases the in-memory secret payload.

Design Details

API Changes

This proposal introduces the following API changes:

  • A new API field on v1.Container: dynamicVolumePolicy (type *ContainerDynamicVolumePolicy), containing restartPolicy (type DynamicVolumeRestartPolicy).
  • Though not introduced by this proposal, we expand the /dynamic subresources introduced in KEP-5972: Dynamic Containers , expanding their schema allowances to permit node-local ephemeral volume mutations.
  • A new API field on v1.PodStatus: volumeStatuses (type []PodVolumeStatus), analogous to containerStatuses, to track whether volumes are allocated and prepared on the host node without coarse pod conditions (actuated container mounts continue to be reported under ContainerStatuses[*].VolumeMounts).

Container Dynamic Volume Policy

To allow users to govern how a container behaves when its volumeMounts are dynamically added or removed (at the same level as resizePolicy in KEP-1287), a new dynamicVolumePolicy field is added to v1.Container:

spec:
  containers:
  - name: runner
    image: registry.k8s.io/pause:3.10
    dynamicVolumePolicy:
      restartPolicy: RestartContainer
    volumeMounts:
    - name: scratch
      mountPath: /mnt/scratch

In go:

type Container struct {
    // ... existing fields ...

    // DynamicVolumePolicy defines dynamic volume mutation behavior for this container.
    // +featureGate=DynamicNodeLocalEphemeralVolumes
    // +optional
    DynamicVolumePolicy *ContainerDynamicVolumePolicy `json:"dynamicVolumePolicy,omitempty"`
}
// ContainerDynamicVolumePolicy defines dynamic volume mutation behavior for a container.
type ContainerDynamicVolumePolicy struct {
    // RestartPolicy defines whether dynamic volume mount additions or removals
    // require restarting the container. Supported values: RestartContainer, "".
    // Defaults to "".
    // +optional
    RestartPolicy DynamicVolumeRestartPolicy `json:"restartPolicy,omitempty"`
}

// DynamicVolumeRestartPolicy defines the restart behavior applied to a container
// when its volume mounts are dynamically mutated.
// +enum
type DynamicVolumeRestartPolicy string

const (
    // DynamicVolumeRestartContainer indicates that Kubelet must restart the container
    // in-place to actuate volume mount additions or removals on a running container.
    DynamicVolumeRestartContainer DynamicVolumeRestartPolicy = "RestartContainer"
)

The semantics of dynamicVolumePolicy.restartPolicy are defined as follows:

  • RestartContainer: Setting dynamicVolumePolicy.restartPolicy: RestartContainer permits adding or removing volume mounts on an existing running container, and explicitly instructs Kubelet to restart (recreate) the container in-place to actuate the mount changes.
  • Default (""): When dynamicVolumePolicy is omitted, the API server rejects requests to add or remove volume mounts while the container is running. Under an empty policy, volume mounts can only be added along with a container addition, and can only be removed along with a container removal.
  • Future (NotRequired): In the future, we will consider supporting hot-plugging volumes to running containers without restart under an additional policy (e.g. NotRequired), but this is out of scope for the first alpha.

Integration with the /dynamic Subresource

KEP-5972: Dynamic Containers introduces the /dynamic subresource. Dynamic volume mutations are executed exclusively via PUT /api/v1/namespaces/{namespace}/pods/{name}/dynamic.

Dynamic volume mutations permit:

  1. Appending new volumes to pod.Spec.Volumes.
  2. Removing volumes from pod.Spec.Volumes if they are no longer referenced by active containers.
  3. Updating volumeMounts within pod.Spec.Containers[*] for containers undergoing restart or addition.

Any semantics regarding the /dynamic subresource that apply to dynamic containers also apply to dynamic volumes.

Inspection via the /allocated Subresource

To provide observability into transactional state and decouple desired state from admitted node state, this proposal integrates with the /allocated subresource introduced by KEP-5972: Dynamic Containers .

The /allocated subresource is served directly by the Kubelet and represents the current transactional allocation on the node. When a pod is updated via /dynamic:

  • pod.Spec.Volumes reflects the desired volume configuration.
  • The /allocated pod endpoint reflects the admitted volume configuration that Kubelet has accepted and is actively preparing or maintaining on the node.
  • The already-existing volumeMounts field under container statuses (pod.Status.ContainerStatuses[*].VolumeMounts) reflects the actual actuated volume mounts in the running container.

Pod Volume Status API (pod.Status.VolumeStatuses)

To track the status of pod volumes across their allocation and host preparation lifecycles, this proposal introduces volumeStatuses to v1.PodStatus. Analogous to containerStatuses, this field provides granular, per-volume observability into whether volumes are allocated and prepared on the host node. Container mount actuation is not tracked in this field, but is surfaced via the preexisting pod.Status.ContainerStatuses[*].VolumeMounts.

The allocated pod, including any dynamically added or removed volumes, is exposed via the already-existing /allocated subresource on Kubelet.

Detailed failure reasons and diagnostics will be emitted via Events, which will be further designed in beta.

type PodStatus struct {
    // ... existing fields ...

    // VolumeStatuses contains the status of volumes in the pod, tracking whether
    // volumes are allocated and prepared on the host node.
    // +listType=map
    // +listMapKey=name
    // +optional
    // +featureGate=DynamicNodeLocalEphemeralVolumes
    VolumeStatuses []PodVolumeStatus `json:"volumeStatuses,omitempty"`
}

// PodVolumeStatus represents the status of an individual volume in a pod.
type PodVolumeStatus struct {
    // Name is the name of the volume matching pod.Spec.Volumes.
    Name string `json:"name"`

    // HostPhase represents the host-level preparation phase of the volume:
    // Pending, HostMounted, ErrorMounting, Unmounting.
    // +optional
    HostPhase PodVolumeHostPhase `json:"hostPhase,omitempty"`

    // ErrorMessage represents the error encountered during the host preparation
    // phase of the volume. It is only populated when an error occurs during SetUpAt.
    // +optional
    ErrorMessage string `json:"errorMessage,omitempty"`
}

// PodVolumeHostPhase represents the host preparation phase of a volume mounted to a pod.
type PodVolumeHostPhase string

const (
    // PodVolumeHostPending indicates that the volume has been added to pod.Spec.Volumes,
    // and is waiting node allocation or host preparation/mounting by VolumeManager.
    PodVolumeHostPending PodVolumeHostPhase = "Pending"

    // PodVolumeHostMounted indicates that VolumeManager has successfully prepared and
    // mounted the volume on the node host (marked ready in ActualStateOfWorld).
    PodVolumeHostMounted PodVolumeHostPhase = "HostMounted"

    // PodVolumeHostErrorMounting indicates that Kubelet admitted the volume into allocatedPod,
    // but VolumeManager encountered an error during SetUpAt while preparing or mounting the
    // volume on the host.
    PodVolumeHostErrorMounting PodVolumeHostPhase = "ErrorMounting"

    // PodVolumeHostUnmounting indicates that volume removal has been admitted and allocated
    // by Kubelet, and Kubelet is actively unmounting container bind mounts and executing host TearDownAt.
    PodVolumeHostUnmounting PodVolumeHostPhase = "Unmounting"
)
Lifecycle Phases
  1. Pending: A dynamic volume mutation was submitted in pod.Spec.Volumes, but Kubelet has not yet admitted it into allocatedPod (for example, due to node ephemeral storage quota limits or a rapid re-add collision). If admission fails, the volume remains in Pending, and Kubelet emits a corresponding event (e.g. FailedAllocation).
  2. HostMounted: Kubelet has admitted the volume into allocatedPod, VolumeManager has successfully executed SetUpAt on the node host, and the volume path (/var/lib/kubelet/pods/<uid>/volumes/...) is recorded ready in ActualStateOfWorld (ASW).
  3. ErrorMounting: Kubelet admitted the volume into allocatedPod, but VolumeManager encountered an error during SetUpAt while preparing or mounting the volume on the host (for example, failure to resolve a referenced Secret or ConfigMap, an image pull failure, or a filesystem mount syscall failure). This ensures failures during the host mounting step are explicitly surfaced rather than leaving the volume in Pending. Detailed error context is recorded in Kubelet Events and in the ErrorMessage field.
  4. Unmounting: The volume removal from pod.Spec.Volumes has been successfully admitted and committed to allocatedPod. Unmounting is set after the removal has been admitted and allocated by Kubelet. If a volume removal is part of a transactional dynamic update that is rejected or deferred during Kubelet allocation, the volume removal is not committed to allocatedPod, and its hostPhase remains HostMounted. Once allocation succeeds, Kubelet retains the volume entry in Unmounting while target containers unmount the volume and VolumeManager executes host TearDownAt. Once host unmount completes and the volume is cleared from ASW, the entry is pruned from VolumeStatuses.
Sample pod.Status Lifecycle States
State 1: Pending Volume and Container

When a dynamic mutation adding a new volume and a container mounting it is submitted via /dynamic, but cannot yet be admitted by Kubelet:

status:
  volumeStatuses:
    - name: dynamic-scratch
      hostPhase: Pending
  containerStatuses:
    - name: worker
      state:
        waiting:
          reason: Unallocated
          message: "Container addition pending node allocation"
State 2: Host Mounted, Container Restart Pending

When a dynamic mutation adding a new volume and mounting it into an existing container is submitted, Kubelet admits the volume into allocatedPod. VolumeManager mounts the host directory in ASW (hostPhase: HostMounted). The following yaml shows the pod status after this step, but before the target container has restarted to pick up the new mount:

status:
  volumeStatuses:
    - name: dynamic-scratch
      hostPhase: HostMounted
  containerStatuses:
    - name: runner
      ready: true
      state:
        running:
          startedAt: "2026-09-25T14:00:00Z"
      # volumeMounts still reflects the pre-restart container mounts
      volumeMounts: []
State 3: Fully Actuated

Kubelet recreates container runner with the volume mount. The container runtime actuates the bind mount into the container namespace, and pod.Status.ContainerStatuses[*].VolumeMounts is updated:

status:
  volumeStatuses:
    - name: dynamic-scratch
      hostPhase: HostMounted
  containerStatuses:
    - name: runner
      ready: true
      state:
        running:
          startedAt: "2026-09-25T14:05:00Z"
      volumeMounts:
        - name: dynamic-scratch
          mountPath: /mnt/scratch
          readOnly: false
State 4: Host Mount Failure (ErrorMounting)

If Kubelet admits the volume into allocatedPod, but VolumeManager encounters an error during SetUpAt on the node:

status:
  volumeStatuses:
    - name: dynamic-scratch
      hostPhase: ErrorMounting
  containerStatuses:
    - name: runner
      ready: true
      state:
        running:
          startedAt: "2026-09-25T14:00:00Z"
      volumeMounts: []

API Server Validation Rules

When a client issues a PUT to /dynamic, the API server validates the request before persisting the pod specification to etcd:

  • Dynamic Volume Policy Validation:
    • Adding or removing volume mounts on an already-running container requires dynamicVolumePolicy.restartPolicy: RestartContainer on that container. If a request attempts to add, remove, or mutate mounts on a running container with an empty policy, the update is rejected with a validation error.
    • Adding mounts with an empty restart policy is permitted only when simultaneously adding a new container (such as a dynamic or ephemeral container).
  • Exclusion of Init Containers (Alpha): Dynamic volume mutations targeting pod.Spec.InitContainers[*] are strictly rejected in Alpha. Any update attempting to add, remove, or mutate volumeMounts on init containers returns a validation error. Supporting init containers (including restartable sidecar init containers) will be reevaluated for Beta.
  • Supported Volume Sources: Only node-local ephemeral volume types are permitted:
    • v1.VolumeSource.ConfigMap
    • v1.VolumeSource.Secret
    • v1.VolumeSource.Projected
    • v1.VolumeSource.EmptyDir
    • v1.VolumeSource.Image
  • Unsupported Types: Updates attempting to add PersistentVolumeClaim, HostPath, CSI, or cloud volume sources are rejected with a validation error.
  • Volume Name Uniqueness and Collision Prevention:
    • API Server Validation: Volume names in spec.volumes must remain unique. In addition, the API server rejects adding a new volume if its name is currently present in pod.Status.VolumeStatuses or still actively reported in any container’s pod.Status.ContainerStatuses[*].VolumeMounts. This prevents race conditions where a client removes a volume and rapidly re-adds one with the same name before container unmounting and host teardown complete, exactly mirroring Dynamic Containers.
  • Dangling Mount Protection: A volume cannot be deleted from spec.volumes if any container in spec.containers still mounts it, unless that container is also being removed or updated to remove the mount in the same transaction.
  • Pod Phase: Dynamic volume mutations are only allowed while the pod is in the Running phase and has completed initialization. Once DeletionTimestamp is set, mutations are rejected.

When pod volumes are updated via /dynamic, the Node Authorizer’s existing informer automatically updates its graph to grant the node access to newly mounted Secrets or ConfigMaps. Any transient propagation delay is safely handled by Kubelet’s volume mounting retry loop.

Two-Stage Kubelet Lifecycle: Allocation and Actuation

Mutations submitted through the /dynamic subresource follow a two-stage Kubelet lifecycle: Allocation and Actuation. While the /dynamic subresource itself comes from KEP-5972: Dynamic Containers , splitting node admission from container runtime configuration adopts the two-stage allocation and actuation model established in KEP-1287: In-Place Pod Vertical Scaling . Once the API server validates the mutation and persists it to etcd, both stages are executed locally on the node by Kubelet.

+-----------------------------------------------------------------------------+
| Stage 1: Allocation (Kubelet Admission & Async Host Preparation)            |
|                                                                             |
| 1. Kubelet admits volume mutation against node-local resources and quotas.  |
| 2. Kubelet checkpoints admitted state to allocatedPod.                      |
| 3. Kubelet reflects admitted volumes in the /allocated pod representation.  |
| 4. VolumeManager DSWP discovers volume in allocatedPod and mounts on host   |
|    asynchronously via SetUpAt in the background.                            |
+-----------------------------------------------------------------------------+
                                       │
                                       ▼
+-----------------------------------------------------------------------------+
| Stage 2: Actuation (Container-Scoped Runtime Configuration)                 |
|                                                                             |
| 1. Initial SyncPod (Pod Startup): Blocks on WaitForAttachAndMount before    |
|    starting any initial containers (preserving standard Kubernetes behavior)|
| 2. Subsequent SyncPods (Dynamic Mutations): Kubelet checks if volumes       |
|    required by target containers are ready in ASW.                          |
| 3. If a dynamic volume is still mounting, only the specific target container|
|    restart/addition is deferred; unaffected containers continue running.    |
| 4. Once ready, Kubelet recreates the target container (or starts the dynamic|
|    container) with bind mounts pointing to the prepared host paths.         |
| 5. Kubelet updates pod.Status.VolumeStatuses (hostPhase) and CRI mounts.   |
+-----------------------------------------------------------------------------+

Stage 1: Admission and Allocation (Node-Local)

When a pod update arrives on the node via watch from the API server:

  1. Kubelet Admission: Kubelet validates that referenced Secrets and ConfigMaps are resolvable and that local ephemeral storage or memory quota is sufficient for requested emptyDir volumes.
  2. State Checkpointing: Kubelet records the admitted volume and container specifications into its local state checkpoint (allocatedPod).
  3. Subresource Exposure: Kubelet reflects the admitted configuration in the /allocated subresource, confirming to external observers that allocation succeeded.
  4. VolumeManager Asynchronous Host Preparation: Kubelet’s VolumeManager is only ever aware of the allocatedPod. The DesiredStateOfWorldPopulator reads volumes exclusively from allocatedPod rather than unadmitted pod specs. This guarantees that VolumeManager never prepares or mounts host directories for mutations that failed local admission. VolumeManager generates the desired state of world (DSW) entry and invokes SetUpAt on the appropriate volume plugin asynchronously in the background.

Stage 2: Actuation (Container Runtime)

Once host allocation and directory preparation are underway, volumes are actuated into target containers during Kubelet’s SyncPod execution:

  1. Initial Pod Startup vs. Subsequent Dynamic SyncPods:
    • Initial Pod Startup: The very first SyncPod invocation for a new pod preserves standard Kubernetes behavior: Kubelet blocks on WaitForAttachAndMount to ensure all creation-time volumes are fully mounted on the host before any containers or init containers start.
    • Subsequent Dynamic SyncPods: For already-running pods processing dynamic mutations, Kubelet avoids globally blocking or stalling the pod worker. Instead, during container reconciliation (computePodActions), Kubelet inspects ActualStateOfWorld (ASW) for the specific volumes required by each container.
  2. Container-Scoped Deferral & Volume Readiness Trigger:
    • Determining Readiness via ASW: Kubelet determines volume readiness by checking VolumeManager’s ActualStateOfWorld (ASW) cache.
    • Container-Scoped Deferral: If a newly added volume required by a container has not yet been marked mounted in ASW, Kubelet defers actuation only for that specific container. Unaffected containers continue running uninterrupted.
    • Event-Driven Trigger: When VolumeManager finishes mounting a volume and marks it in ASW, it triggers a pod sync to immediately actuate the deferred container.
  3. Explicit Container Re-creation Decisions: When the required volumes are verified ready in ASW, computePodActions detects that an existing container’s volumeMounts in allocatedPod differ from the running container configuration and schedules an explicit container restart (stop and recreate). For newly introduced dynamic or ephemeral containers, Kubelet schedules container creation.
  4. CRI Invocation: Kubelet stops the existing container if restarting, constructs the CRI container specification with bind mounts pointing to the prepared host paths, and invokes CRI CreateContainer followed by StartContainer.
  5. Status and Actuation Reflection: Kubelet updates pod.Status.VolumeStatuses to reflect host-level volume readiness (hostPhase: HostMounted, or ErrorMounting on failure). Actuation into containers is reflected exclusively under pod.Status.ContainerStatuses[*].VolumeMounts as CRI confirms the active bind mounts. User observability into allocation details and failure causes is provided via Kubelet Events and the /allocated subresource rather than bloating PodStatus.

Volume Teardown and Removal Lifecycle

Volume removal follows the same two-stage allocation and actuation pattern, ensuring that container unmounting precedes host directory teardown:

  1. API Validation: The API server strictly rejects removing a volume from spec.volumes if any container in spec.containers still mounts it, unless that container mount is also removed in the same /dynamic update.
  2. Allocation Stage (Atomic Admission): When a /dynamic update removes a volume along with its referencing container (or removes the mount from an existing container), Kubelet admits the mutation and commits the changes to allocatedPod atomically.
  3. Actuation Stage (Container Unmount): During actuation, Kubelet’s SyncPod acts on the container first, stopping and recreating it with the updated container specification (excluding the removed volume mount). The container runtime releases the mount from the container’s mount namespace.
  4. Teardown & Host Unmount via PodStateProvider: VolumeManager’s DesiredStateOfWorldPopulator (DSWP) observes that the volume has been removed from allocatedPod. To prevent tearing down host directories while containers are still running, DSWP queries Kubelet via an extension to the PodStateProvider interface:
    type PodStateProvider interface {
        // IsVolumeInUseByPod returns true if any running container in the pod
        // currently mounts the specified volume.
        IsVolumeInUseByPod(podUID types.UID, volumeName string) bool
    }
    
    Kubelet implements IsVolumeInUseByPod by checking its internal, in-memory rutime cache of running container statuses. Once CRI confirms the old container has stopped, the runtime cache clears the mount reference. Only then does DSWP remove the volume from DesiredStateOfWorld (DSW). The reconciler then unmounts and remove the host directory.
  5. Status Pruning: When a volume removal is submitted via /dynamic, its entry in pod.Status.VolumeStatuses transitions to hostPhase: Unmounting only after the removal has been successfully allocated by Kubelet. If the removal is part of a transactional dynamic update that is rejected or deferred during Kubelet allocation (for example, if submitted alongside an infeasible container resource resize), the volume removal is not committed to allocatedPod, and its hostPhase remains HostMounted. Once allocated, Kubelet retains Unmounting while containers are restarted or stopped without the volume mount (updating pod.Status.ContainerStatuses[*].VolumeMounts). Once PodStateProvider.IsVolumeInUseByPod returns false and VolumeManager completes host TearDownAt (clearing ASW), Kubelet prunes the volume entry completely from pod.Status.VolumeStatuses.

Limitations

  • Actuation occurs only upon container restart, addition (ephemeral or dynamic containers), or removal.
  • Live hot-plug into running containers without restart (NotRequired) is deferred to a future milestone or enhancement.
  • Dynamic volume mutations for initContainers (including restartable sidecar init containers) are not supported in Alpha and will be reevaluated for Beta.
  • Dynamic addition of hostPath volumes is not supported in Alpha due to security considerations and will be reevaluated for Beta.

Test Plan

[x] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.

Prerequisite testing updates

None.

Unit tests
  • API validation:
    • Adding supported node-local volume types (configMap, secret, emptyDir, projected, image) via /dynamic.
    • Rejection of unsupported types (PVC, hostPath, CSI).
    • Rejection of duplicate volume names and dangling container mounts.
  • Node Authorizer:
    • Graph edge creation for dynamically mounted Secrets and ConfigMaps on running pods.
  • Kubelet VolumeManager populator:
    • Discovery of newly added volumes from allocatedPod on running pods.
    • Teardown of removed volumes while containers remain running.
  • Volume plugins (configmap, secret, emptydir, projected):
    • Setup and teardown lifecycle on dynamic add/remove.
Integration tests
  • API Server /dynamic subresource endpoint:
    • RBAC enforcement (cluster-admin required).
    • Updates to spec.volumes and volumeMounts succeed.
  • Node Authorizer integration:
    • Kubelet can retrieve Secrets mounted dynamically into a running pod.
e2e tests
  • Dynamic emptyDir on container restart:
    • Start a pod, dynamically add an emptyDir volume, restart the container in-place, and verify the container writes data to the new mount.
  • Dynamic ConfigMap and Secret addition:
    • Add a ConfigMap and Secret volume dynamically, restart the container, and verify keys are mapped to files inside the container.
  • Dynamic volume removal:
    • Remove a volume from the pod spec, verify the host directory is unmounted and cleaned up by Kubelet.
  • Synergy with KEP-5972: Dynamic Containers :
    • Add a dynamic container and a new volume simultaneously via /dynamic; verify container starts successfully with the volume mounted.

Graduation Criteria

Alpha

  • Feature gate DynamicNodeLocalEphemeralVolumes implemented (default disabled).
  • dynamicVolumePolicy field added to v1.Container (with nested restartPolicy).
  • Support for node-local volumes (configMap, secret, projected, emptyDir, image) via /dynamic.
  • Allocation and actuation on container restart, container addition, and container removal.
  • Unit, integration, and e2e tests passing.

Beta

  • Consider enabling live hot-plug into running containers, relaxing validation on container.dynamicVolumePolicy.restartPolicy to permit NotRequired for live mutations. This may be added in alpha2 or a separate KEP.
  • Reevaluate supporting hostPath volumes with appropriate security safeguards.
  • Reevaluate dynamic volume mutations for initContainers (including restartable sidecar init containers).
  • Explore the feasibility of allowing Kubelet to retry admitting a rapidly re-added volume after the previous volume has completed host unmount and cleared ASW.
  • Metrics for dynamic volume mount and unmount latency and error rates.
  • Evaluate integration with pod.status.volumeHealth (KEP-1432: Volume Health Monitor) as CSIVolumeHealth stabilizes, particularly for reporting underlying volume faults alongside dynamic volume lifecycle states.

GA

  • Allowing time for feedback, with at least 2 release cycles in Beta / enabled by default.
  • Further GA criteria to be added in the beta update.

Upgrade / Downgrade Strategy

  • Upgrade:
    • Enabling the feature gate introduces the /dynamic volume capability. Existing pods are unaffected.
  • Downgrade:
    • Disabling the feature gate restores strict immutability. Pods that previously added volumes continue running with their current mounts, but further /dynamic updates to volumes are rejected.

Version Skew Strategy

Version skew will be handled by NodeDeclaredFeatures. If a client issues an update request to a Node that does not support it, the API server will reject the request.

Production Readiness Review Questionnaire

Feature Enablement and Rollback

How can this feature be enabled / disabled in a live cluster?
  • Feature Gate: DynamicNodeLocalEphemeralVolumes
  • Components: kube-apiserver, kubelet
Does enabling the feature change any default behavior?

Standard updates to /pods remain immutable, but third-party controllers and ecosystem tools may assume that a pod’s volumes field is stable and immutable after creation. Enabling this feature changes that assumption, as spec.volumes can now be mutated dynamically via the /dynamic subresource.

The /dynamic subresource was intended to capture the intent that “all fields may be made mutable under this subresource”, so using it for this purpose should not be surprising to users.

Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?

Yes. Disabling the feature gate prevents any further dynamic volume mutations via the API.

What happens if we reenable the feature if it was previously rolled back?

The API server resumes accepting dynamic volume updates on /dynamic.

Are there any tests for feature enablement/disablement?

Yes. Unit and integration tests will verify that disabling the feature gate rejects dynamic volume updates with a feature disabled validation error.

Rollout, Upgrade and Rollback Planning

How can a rollout or rollback fail? Can it impact already running workloads?
What specific metrics should inform a rollback?
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?

Monitoring Requirements

How can an operator determine if the feature is in use by workloads?
How can someone using this feature know that it is working for their instance?
  • Events
    • Event Reason:
  • API .status
    • Condition name:
    • Other field:
  • Other (treat as last resort)
    • Details:
What are the reasonable SLOs (Service Level Objectives) for the enhancement?
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
  • Metrics
    • Metric name:
    • [Optional] Aggregation method:
    • Components exposing the metric:
  • Other (treat as last resort)
    • Details:
Are there any missing metrics that would be useful to have to improve observability of this feature?

Dependencies

Does this feature depend on any specific services running in the cluster?

Scalability

Will enabling / using this feature result in any new API calls?

Yes, updates to /dynamic generate write requests and pod update watch events.

Will enabling / using this feature result in introducing new API types?

No, there are new API types.

Will enabling / using this feature result in any new calls to the cloud provider?

No. Node-local volumes have no cloud provider interactions.

Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?

No. Dynamic volume mutation is a new operation not covered by existing SLIs/SLOs.

Will enabling / using this feature result in increasing size or count of the existing API objects?

Modest increase in the size of v1.Pod objects when additional volume entries are declared or when the volume restart policy is set.

Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?

Negligible CPU and memory impact on Kubelet. Disk usage corresponds to files created in emptyDir.

Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?

No. Dynamic volumes do not affect this kind of node resources.

Troubleshooting

How does this feature react if the API server and/or etcd is unavailable?
What are other known failure modes?
What steps should be taken if SLOs are not being met to determine the problem?

Implementation History

  • 2026-09-21: KEP created for Alpha.

Drawbacks

  • Allowing .spec.volumes to mutate departs from Kubernetes’’’ historical assumption of static pod specs, requiring ecosystem controllers to adjust their reconciliation models.

Alternatives

Alternative 1: Out-of-band HostPath Bind Mounts

Workloads mount a shared host directory and create subdirectories manually.

  • Why rejected: Bypasses Kubernetes RBAC, volume authorization, quotas, and security boundaries.

Alternative 2: Full Pod Recreation

Destroy and recreate the pod whenever storage requirements change.

  • Why rejected: Unacceptable latency (seconds to minutes) for high-frequency agentic and sandboxed workloads.

Alternative 3: Dedicated Dynamic Volume Container Type

Introduce a new volume type (e.g. dynamicVolumePool) that contains a mutable list of nested volumes.

  • Why rejected: Adds significant API surface complexity. Mutating spec.volumes directly via /dynamic provides a cleaner, uniform experience that naturally mirrors container dynamism in KEP-5972: Dynamic Containers .

Alternative 4: Live Hot-Plug into Running Containers for Alpha

Support live in-place hot-mounting into running containers without restart in Alpha via mount propagation.

  • Why rejected for Alpha: While technically feasible, it adds operational complexity (requiring parent mount setup and specific directory conventions). Deferring to Beta allows the core API and VolumeManager lifecycle to stabilize first.