KEP-3953: In-place Node Resource Resize
KEP-3953: In-place Node Resource Resize
- Release Signoff Checklist
- Glossary
- Summary
- Motivation
- Proposal
- Layer 1: Ecosystem Tolerance for Mutable Node Capacity
- Layer 2: Declarative Capacity Actuation via
Node.Spec.ConfiguredCapacity(targeted for v1.39 Alpha) - User Stories
- Story 1: Maximizing Specialized Hardware
- Story 2: Vertical Scaling for Performance
- Story 3: Reducing Operational Complexity (Scale-Up vs. Scale-Out)
- Story 4: Instant Capacity Utilization
- Story 5: Zero-Disruption Operations
- Story 6: Emergency Hardware Removal (Self-Protection)
- Story 7: Dynamic Storage Expansion
- Story 8: Orchestrated Capacity Downscale
- Notes/Constraints/Caveats (Optional)
- Risks and Mitigations
- Design Details
- API Changes
- Node Conditions for State Dissemination
- Admission Control Contract
- Resource-Specific Validation Rules
- Security Considerations
- Architecture Flow
- Path A1: Orchestrated Configuration-Driven Upscale
- Path A2: Orchestrated Configuration-Driven Downscale
- Path B: Emergency Hardware-Driven Fallback (Ungraceful Downscale)
- Flow Control: Container Swap Limit Recalculation
- Flow Control: Hardware Degradation and Capacity Starvation
- Compatibility with Cluster Autoscaler
- Layer 1 Implementation: Ecosystem Tolerance — Validating That Mutable Capacity Is Safe
- Layer 2 Implementation: Declarative Capacity Actuation
- Observability and Metrics
- Test Plan
- Graduation Criteria
- Upgrade / Downgrade Strategy
- Version Skew Strategy
- Production Readiness Review Questionnaire
- Implementation History
- Drawbacks
- Alternatives
- Open Questions for Layer 2
- Infrastructure Needed (Optional)
- Future Work
Release Signoff Checklist
Items marked with (R) are required prior to targeting to a milestone / release.
- (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
- (R) KEP approvers have approved the KEP status as
implementable - (R) Design details are appropriately documented
- (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
- e2e Tests for all Beta API Operations (endpoints)
- (R) Ensure GA e2e tests meet requirements for Conformance Tests
- (R) Minimum Two Week Window for GA e2e tests to prove flake free
- (R) Graduation criteria is in place
- (R) all GA Endpoints must be hit by Conformance Tests
- (R) Production readiness review completed
- (R) Production readiness review approved
- “Implementation History” section is up-to-date for milestone
- User-facing documentation has been created in kubernetes/website , for publication to kubernetes.io
- Supporting documentation—e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes
Glossary
- In-Place Resource Resize: Dynamically increasing or decreasing compute resources (CPU, Memory, and HugePages) on a live,
Readynode via a declarative API, with the Kubelet reconciling the change in place. Swap is explicitly out of scope for Alpha — see Non-Goals. - Node Compute Resource: CPU, Memory, and HugePages. Swap is deferred to a later milestone.
- Physical Capacity: The raw hardware capacity of a node as reported by the integrated
cAdvisorsubsystem, reflecting the true underlying machine resources (e.g., number of CPU cores). - Configured Capacity: The desired logical capacity declared by an administrator or external controller via
Node.Spec.ConfiguredCapacity. This is the Kubelet’s target and may be less than the Physical Capacity (e.g., to under-report resources intentionally). For Alpha, Configured Capacity must not exceed Physical Capacity. - CapacityConfigured Condition: A
Node.Status.Conditionof typeCapacityConfiguredthat the Kubelet uses to expose the current reconciliation state of a capacity resize request. Possible reasons areAccepted,InProgress,Infeasible, andEmergencyReduced. - Kubelet-Restart Workaround: The pre-existing, informal practice of restarting the Kubelet after a hardware change so that it re-reads physical capacity from cAdvisor at boot time. This is not a supported or safe mechanism for mutating node capacity — it is a workaround with significant drawbacks (downtime, edge-case bugs, no ecosystem coordination).
Summary
This proposal facilitates dynamic native resource resizing (increases and decreases in capacity) on a node to streamline cluster capacity updates, offering a seamless alternative to adding or removing nodes from an existing cluster. The revised node configurations automatically propagate at both the node and cluster levels.
Today, mutating Node.Status.Capacity is not a supported or safe operation in Kubernetes. The Kubernetes ecosystem (Scheduler, Cluster Autoscaler, VPA, Quota system) has never been explicitly designed or tested to tolerate a live change to a Node’s capacity values. This KEP addresses the problem in two ordered layers:
Layer 1 — Ecosystem Tolerance: Formally define the contract that the broader cluster ecosystem must honour when Node.Status.Capacity changes. Establish the API-server mutations, Scheduler cache-invalidation tests, Cluster Autoscaler behaviour validation, and Quota/VPA integration points that make a capacity change safe to propagate — even if, as a first step, that change is initiated by a Kubelet restart. This is the scope of the v1.38 Alpha milestone.
Layer 2 — Declarative Capacity Actuation: Introduce Node.Spec.ConfiguredCapacity, the declarative API field that allows external controllers to trigger a capacity change. Implement the Kubelet reconciliation loop that updates cgroups, re-initialises sub-managers (CPU Manager, Memory Manager), recalculates container swap limits via the CRI, and synchronises Node.Status — all on the running node. The CapacityConfigured Node Condition disseminates the live reconciliation state to external orchestrators. This layer is fully described in this KEP and targets a subsequent milestone (v1.39 Alpha or later), pending resolution of the open design questions documented in Open Questions for Layer 2
.
The full design of Layer 2 — including the Node.Spec.ConfiguredCapacity API field, the Kubelet reconciliation loop, and the CapacityConfigured condition — is described in this document so that the community can evaluate the complete architecture. However, no Layer 2 code is introduced in v1.38; the initial milestone is deliberately scoped to proving the ecosystem foundation is safe, using the existing Kubelet-restart-on-resized-hardware path as the trigger.
Motivation
The Problem: Live Mutation of Node Capacity Is Unsupported Today
Performing a live mutation of Node.Status.Capacity on a running node is not a currently supported operation in Kubernetes. A node’s resource capacity values (Node.Status.Capacity, Node.Status.Allocatable) are written by the Kubelet during bootstrap and then held fixed for the entire lifetime of that Kubelet process. When the Kubelet restarts, it re-reads hardware capacity from cAdvisor and rewrites these fields — so a restart does cause the values to change. However, this is a full process restart that reinitialises the Kubelet from scratch, not a live mutation of a running node’s capacity. No existing API validation, Scheduler logic, Cluster Autoscaler heuristic, VPA algorithm, or ResourceQuota controller has been designed or tested against the assumption that these fields can change for a live, Ready node without a Kubelet restart.
The only known workaround today is to restart the Kubelet after a physical hardware change. However, this workaround:
- Does not safely mutate capacity on a live node — the Kubelet reinitialises entirely from scratch with no coordination with the Scheduler, Cluster Autoscaler, or Quota system before or after the change.
- Is not a first-class supported operation — there are no established best practices for coupling a Kubelet restart to a cloud provider or hypervisor API call.
- Carries significant operational risk — it introduces temporary node
NotReadystates, breaks activeexec/port-forwardsessions, and has a documented history of triggering edge-case bugs:
In summary: a Kubelet restart is a workaround for the absence of a supported capacity-mutation path, not a safe or reliable mechanism for it. This KEP introduces that supported path.
The Need: Dynamic Hardware Requires a Dynamic Capacity Model
Modern hypervisors, cloud providers, and kernel capabilities enable the dynamic hot-plugging and hot-unplugging of native resources such as CPU and Memory (e.g., CPU Hotplug , Memory Hotplug ), and Ephemeral Storage block devices. Without a supported capacity-mutation path, these hardware events lead to two severe failure modes:
Cluster-Level Starvation (Upscaling): If capacity is added, the Kubernetes Scheduler and Cluster Autoscaler remain blind to it, rendering the new hardware useless for pending workloads.
Node-Level Instability (Downscaling): If capacity is removed, the Kubelet’s stale top-level cgroups, container swap limits, and Eviction Manager thresholds do not adjust. Because the Kubelet assumes the resources still exist, it fails to evict pods defensively, causing the host’s Linux kernel to invoke the OOM killer or exhaust the disk, violently terminating processes and potentially crashing the node.
The Two-Layer Solution
This KEP decomposes the solution into two ordered, independently valuable layers:
Layer 1: Making the Ecosystem Safe for Mutable Capacity
Before any Kubelet actuation can be trusted, the broader cluster ecosystem must be formally validated and — where necessary — updated to tolerate Node.Status.Capacity mutations safely. This layer uses the Kubelet restart as its initial boundary: it answers the question “If the capacity field changes, even as a result of a Kubelet restart on resized hardware, are the API Server, Scheduler, Cluster Autoscaler, VPA, and Quota system safe?”
The ecosystem contracts that must be established:
- API Server:
Node.Status.CapacityandNode.Status.Allocatablemust be patchable on a liveReadynode without triggering systemic webhook rejections or validation failures. This behaviour has never been formally tested or documented. - Scheduler: The Scheduler’s internal node cache must correctly invalidate and update its per-node resource view when it receives a
NodeUPDATE event with changed capacity fields. It must not rely on a stale boot-time snapshot. - Cluster Autoscaler (CA): The CA uses existing nodes as templates for provisioning new nodes within the same NodeGroup. If capacity is mutable, the CA must be given a stable, immutable baseline (via a boot-time annotation) to template from, rather than reading the live, potentially-resized
Node.Status.Capacity. - VPA (Vertical Pod Autoscaler): The VPA’s node capacity model must be tolerant of live capacity changes so that it does not issue recommendations that exceed the new node bounds.
- ResourceQuota: Namespace-level ResourceQuota controllers operate against allocatable node capacity. A capacity change must not silently bypass quota enforcement.
Layer 2: Declarative Capacity Actuation
Once the ecosystem is safe (Layer 1), the second layer provides the supported, declarative mechanism for capacity actuation on a running node. This layer introduces:
Node.Spec.ConfiguredCapacity: A new declarative API field allowing external controllers or administrators to declare the desired logical capacity without touching the node or its Kubelet process.- Kubelet Reconciliation Loop: The Kubelet watches for changes to
Node.Spec.ConfiguredCapacityand for physical hardware drift detected by cAdvisor. It validates the desired value against physical bounds, then actuates the change by updating internal cgroup hierarchies, re-initialising sub-managers (CPU Manager, Memory Manager, Eviction Manager), and recalculating container swap limits via the CRI — all in a running Kubelet. CapacityConfiguredNode Condition: A new condition that exposes the live reconciliation state (Accepted,InProgress,Infeasible,EmergencyReduced) to external orchestrators, enabling safe, signal-driven coordination for downscale workflows.
Why a Declarative API Is Essential for Safe Downscaling
Without an API-driven trigger, cluster administrators have no safe way to coordinate a hot-unplug operation. The desired workflow for a graceful downscale is:
- An external controller invokes the API (
Node.Spec.ConfiguredCapacity) to reduce logical node capacity. - The Kubelet actuates that reduction (graceful eviction, updates cgroups, updates
Node.Status). - The external controller observes the Kubelet’s completion signal (
CapacityConfigured: Accepted). - The external controller proceeds to physically remove hardware via the hypervisor.
This coordination is impossible with a purely reactive, hardware-first model. If the hypervisor forcefully reclaims RAM (e.g., balloon deflation) without prior API coordination, the kernel OOM killer may fire before the Kubelet can react. By making the API the primary trigger, operators gain deterministic control over the timing and safety of capacity removal.
Benefits Beyond the Workaround
Enabling the Kubelet to dynamically detect and adapt to underlying capacity changes mitigates manual administrative toil and unlocks several distinct advantages over the current Kubelet-restart workaround:
- Control Plane Efficiency: Managing resource demands by scaling existing nodes in-place brings significantly less overhead to the control plane compared to provisioning and joining entirely new nodes.
- Speed to Delivery: Expanding the capabilities of current nodes is considerably more time-efficient than the procedure of establishing new virtual machines.
- Network Optimization: Improved inter-pod network latencies, as inter-node traffic is reduced when more pods can be hosted locally on a single scaled-up node.
- Stability: Avoids the historical bugs and disruption risks associated with forced Kubelet restarts.
Implementing this KEP will empower nodes to recognize and adapt to changes in their native configurations instantly, facilitating the safe, efficient, and uninterrupted deployment of workloads.
Goals
API Synchronization: Update Node API Capacity and Allocatable fields dynamically on a live,
Readynode via the declarativeNode.Spec.ConfiguredCapacityAPI field.Component Sync: Re-initialize internal Kubelet managers (CPU, Memory, Eviction) to safely align with the altered hardware capacity.
Cgroup Enforcement: Update the host’s top-level /kubepods and QoS cgroup boundaries to physically enforce the resized limits.
Configured Capacity: Allow the logical capacity of a node to be dynamically configured via
Node.Spec.ConfiguredCapacity, decoupling the cluster’s view of the node from strict physical hardware events. Both hardware-triggered and purely configuration-driven changes (e.g., under-reporting a 32Gi machine as 20Gi) are in-scope.Bootstrap Parity: Upon Kubelet restart, the Kubelet reads
Node.Spec.ConfiguredCapacityas its primary target before falling back to raw cAdvisor hardware discovery. This ensures that a Kubelet restart on a node with an existingConfiguredCapacityspec behaves identically to a live-resize event.
Non-Goals
Reserved Adjustments: Dynamically changing –system-reserved and –kube-reserved values (these remain static from bootstrap).
Infrastructure Orchestration: Executing the physical hardware hot-plug or updating the autoscaler to trigger it.
Workload Re-balancing: Automatically migrating or re-balancing existing workloads across the cluster to utilize the new space.
NRI Plugins: Propagating host resource changes to external Node Resource Interface (NRI) plugins.
OOM Score Updates: Dynamically rewriting oom_score_adj for running processes, due to severe latency and race condition risks.
Pod Resizing: Dynamically resizing individual Pod resource requests and limits (covered independently by KEP-1287).
Swap-Enabled Nodes: In-place node resize is not supported on nodes with Swap enabled for the Alpha phase. The interaction between node capacity resize and per-container swap limit recalculation is non-trivial, and there is ongoing work to align the swap semantics between node resize and pod resize (KEP-1287). To avoid compounding those open questions, resize operations on swap-enabled nodes will be rejected or skipped until a later milestone. This will be revisited in a subsequent Alpha or Beta update.
Node Capacity Overcommit: Configuring the Kubelet to report a logical capacity to the API Server that exceeds the raw, physical underlying hardware capacity (e.g., reporting 48Gi on a 32Gi machine relying on swap). For the Alpha phase,
ConfiguredCapacityis strictly bounded by physical reality (CPU and Memory). This will be explored in Future Work.Admission Webhook on Node.Spec: This KEP does not introduce a new admission webhook specifically for capacity changes. Standard Kubernetes
ValidatingWebhookConfigurationandMutatingWebhookConfigurationcan be deployed by cluster administrators to intercept mutations toNode.Spec.ConfiguredCapacitywithout any KEP-specific mechanism.
Proposal
This KEP introduces a declarative, event-driven reconciliation architecture to handle native resource reconfiguration safely. It is structured as two ordered layers that build on each other, with each layer being independently valuable and reviewable.
Layer 1: Ecosystem Tolerance for Mutable Node Capacity
This layer addresses the foundational question that has never been formally answered: can the Kubernetes control plane safely tolerate a change to Node.Status.Capacity? Regardless of whether that change originates from a Kubelet restart on resized hardware, a manual patch, or the declarative reconciliation loop introduced in Layer 2, every downstream component must behave correctly.
The requirements for Layer 1 are:
API Server — Capacity Field Mutability: Formally validate that
Node.Status.CapacityandNode.Status.Allocatableare patchable on a live,ReadyNode object. Confirm that no system webhook, validation rule, or strategy-merge logic prevents this mutation. This is the unspoken prerequisite that all subsequent work depends on.Scheduler — Node Cache Invalidation: Confirm that the Scheduler’s internal
NodeInfocache is invalidated and refreshed when it receives aNodeUPDATE event with changed capacity fields. The Scheduler must not rely on a snapshot taken at node registration time. This is validated by: (a) upscaling a node’s capacity and verifying a previously unschedulable pod becomes schedulable, and (b) downscaling a node’s capacity and verifying the Scheduler correctly rejects pods that no longer fit.Scheduler Capacity View During Downscale: There is an inherent race condition where the Scheduler schedules a pod to a node concurrently with an in-progress downscale: the Scheduler’s view of the node may still reflect the pre-downscale capacity, but by the time the pod reaches Kubelet admission, the Kubelet’s capacity has already been reduced — causing Kubelet to reject the pod and leave it in a
Failedstate. This is not a new problem (it is structurally identical to a node goingNotReadybetween scheduling and binding). The mitigation — having the Scheduler treat the effective node capacity asmin(Node.Spec.ConfiguredCapacity, Node.Status.Capacity)— is a Layer 2 change: it requiresNode.Spec.ConfiguredCapacityto exist, which is only introduced in Layer 2. Using the minimum means that as soon as an operator signals a downscale viaNode.Spec.ConfiguredCapacity, the Scheduler conservatively stops over-committing to the node even before the Kubelet has finished reconcilingNode.Status. This mirrors the analogous treatment for in-place pod resize, where the Scheduler usesmax(pod.spec.resources, pod.status.resources)to avoid over-committing against a pod whose resources are still being expanded. This strategy minimises the race window but does not eliminate it entirely; Kubelet admission remains the authoritative gate and the final safety net. The Layer 1 deliverable for this item is to document the race and validate that Kubelet admission correctly rejects the over-committed pod — themin()mitigation is implemented as part of Layer 2.Preemption Grace Period Race: A more specific variant of the above concerns the Scheduler’s preemption path. When a high-priority pod arrives and the Scheduler determines it fits on a node only after preempting a lower-priority victim, the victim enters its termination grace period. If node capacity decreases during that grace period, the Scheduler’s original preemption calculation becomes stale: the capacity it assumed the high-priority pod would land on no longer exists.
The Scheduler currently has no general mechanism to detect that a previously computed preemption is no longer sufficient and re-run the calculation. Addressing this in full is planned for the v1.39 cycle.
However, one narrower case can be handled today: if a pod has been nominated to a node (i.e. preemption was selected and the pod is waiting for the victim to terminate) and a Node UPDATE event reduces
Node.Status.Allocatablebelow what the nominated pod requires, the Scheduler can and should re-queue that pod immediately to trigger a fresh scheduling cycle. The Layer 1 deliverable for this item is to add or validate a Node UPDATE queueing hint that re-queues pods nominated to a node whenNode.Status.Allocatableon that node decreases. Without this hint, the nominated pod may attempt to bind to a node that can no longer accommodate it — falling back to Kubelet admission rejection — or may wait indefinitely for space that cannot materialise.Cluster Autoscaler — Stable Provisioning Template: The CA currently uses an existing node’s
Node.Status.Capacityas the template for provisioning new nodes in the same NodeGroup. With mutable capacity, a dynamically resized node must not corrupt this template. The Layer 1 deliverable for this item is to document the current CA behaviour and identify the gap — not to ship a fix. The correct fix is an open design question: placing a static boot-time value on the Node object (whether as an annotation or a new status field) is problematic becauseNode.Statusis meant to reflect live state, annotations are brittle when multiple actors read them, and the edge case where an entire NodeGroup has been uniformly resized means the “initial” capacity is no longer functionally accurate. Alternative approaches — such as having CA read directly from the cloud provider’s launch template, or introducing a configurable reference capacity within CA’s own NodeGroup configuration — are being considered. This question is tracked as an open design item and must be resolved before this KEP’s CA integration is considered complete.VPA — Recommendation Bounds: Confirm that the Vertical Pod Autoscaler’s node-capacity model is re-read on Node UPDATE events. VPA must not issue container resource recommendations that exceed the new node allocatable bounds.
ResourceQuota — Allocatable Accounting: Confirm that namespace-level ResourceQuota admission does not cache the node’s allocatable value at pod admission time in a way that silently bypasses quota limits after a capacity change.
In-Place Pod Resize Interaction (KEP-1287): Node capacity changes interact directly with the
DeferredandInfeasibleresize states defined by In-Place Pod Vertical Scaling (KEP-1287). Three cases must be explicitly handled:Upscale → retry
Deferredpod resizes: WhenNode.Status.Allocatableincreases, pod resize requests that were previously markedDeferred(because insufficient node capacity prevented the Kubelet from accepting them) must be re-evaluated. The Kubelet should attempt the deferred resize against the new, larger allocatable values. This mirrors the existing behaviour where a pod resize that isDeferreddue to resource contention is retried when other pods are evicted and space becomes available — a node capacity upscale is the same signal at a different granularity.Downscale →
DeferredbecomesInfeasible: If a pod resize was previouslyDeferred(waiting for capacity to become available) and a subsequent node downscale reduces allocatable resources to a point where the desired pod resources can never fit, the resize status must transition fromDeferredtoInfeasible. Leaving a resize inDeferredstate against a node that is now provably too small is misleading and prevents the Scheduler and VPA from taking corrective action.Infeasiblepod resize cleared by upscale: Since Kubernetes 1.36, the API server rejects pod resize requests that exceed node capacity at admission. However, the Kubelet can independently mark a resizeInfeasiblewhen capacity-bound constraints are evaluated at the node level. WhenNode.Status.Allocatablesubsequently increases (via a node upscale), previouslyInfeasiblepod resize requests that were blocked solely due to node capacity must be re-evaluated and, if now feasible, transitioned out of theInfeasiblestate. Without this re-evaluation, a node upscale would not unblock any waiting workloads. The Kubelet needs a mechanism to track the reason a resize was markedInfeasible(node-capacity-bound vs. other reasons) in order to know which requests to retry.
The Scheduler and VPA both react to node capacity constraints and pod resize statuses in different ways. The Scheduler operates on
Node.Status.Allocatableand pod resource requests through its standard scheduling cycle: a node capacity change that reduces allocatable resources may cause it to preempt lower-priority pods when higher-priority pods need to be placed — preemption is driven by capacity, not by resize status directly. VPA, by contrast, explicitly observes resize outcomes and may lower its recommendation if a resize is persistentlyInfeasible. A node capacity change can therefore trigger a cascade across all three components — node resize → pod resize state update → Scheduler preemption evaluation and VPA re-evaluation. This interaction chain must be documented and validated as part of the Layer 1 ecosystem contracts.
Layer 1 establishes the baseline: documenting the current ecosystem behaviour when Node.Status.Capacity changes and identifying any gaps. Where gaps exist, they are addressed as part of this KEP — either by fixing the relevant component or by adding the missing contract. Layer 2 then builds the supported actuation mechanism on top of that validated foundation.
Layer 2: Declarative Capacity Actuation via Node.Spec.ConfiguredCapacity (targeted for v1.39 Alpha)
Scope note: Layer 2 is described here for completeness and community review. It is not part of the v1.38 Alpha scope. Implementation begins once the Layer 1 ecosystem contracts are validated and the open design questions in Open Questions for Layer 2 are resolved.
With the ecosystem contracts from Layer 1 established, Layer 2 provides the supported, first-class mechanism for capacity actuation. It is structured as three implementation steps that build incrementally:
Step 1 — Baseline API & Scheduler Validation (pre-requisite tests): Add the upstream test coverage described in Layer 1 before any dynamic trigger code is written. These tests form the safety net for all subsequent changes.
Step 2 — Unified Kubelet Reconciliation: Introduce Node.Spec.ConfiguredCapacity and update the Kubelet to reconcile the difference between what Node.Status reflects and what the Kubelet observes from cAdvisor — covering both the live-node case and the Kubelet-restart-on-resized-hardware case. The Kubelet updates internal cgroup hierarchies, re-initialises sub-managers (CPU Manager, Memory Manager, Eviction Manager), and recalculates container swap limits via the CRI. Node.Status.Capacity is the output of this reconciliation — written by the Kubelet after validation, not used as an input.
Step 3 — Hardware-Drift Trigger: Implement an automated cAdvisor polling trigger that fires the reconciliation loop when physical capacity changes. The loop re-validates Node.Spec.ConfiguredCapacity against the new physical bounds and, only upon successful validation, writes the resolved capacity to Node.Status. This step extends Step 2 to handle emergency hardware-removal events (Path B) where no prior API signal was sent.
By separating these two layers, the KEP ensures that the requirements for making capacity mutable at the cluster level are addressed before the declarative reconciliation mechanism is built on top of them.
User Stories
Story 1: Maximizing Specialized Hardware
As a Cluster Administrator, I want to seamlessly add CPU and memory to an existing node equipped with specialized, scarce hardware (e.g., custom ASICs, specific CPU architectures), so that I can maximize the utilization of that specific hardware without draining the node or disrupting the workloads already utilizing it.
Story 2: Vertical Scaling for Performance
As a Performance Engineer, I want to dynamically increase the compute capacity of a node running a monolithic database or data-heavy application, so that the application can immediately benefit from larger memory caches and reduced context-switching without suffering the downtime of a pod migration.
Story 3: Reducing Operational Complexity (Scale-Up vs. Scale-Out)
As a Site Reliability Engineer (SRE), I want the option to vertically scale existing nodes instead of always horizontally provisioning new VMs, so that I can manage fewer, larger nodes to simplify network topology, monitoring overhead, and overall cluster complexity.
Story 4: Instant Capacity Utilization
As a Cluster Administrator, I want the Kubernetes control plane to instantly recognize when a node’s capacity is expanded on-the-fly via my cloud provider, so that pending workloads can be scheduled onto that new space immediately without the latency of waiting for a new VM to boot and join the cluster.
Story 5: Zero-Disruption Operations
As an Application Owner, I expect my running workloads to experience zero downtime or disruption when the infrastructure administrator adds capacity to the underlying node, entirely avoiding the historical risks and bugs associated with forced Kubelet restarts or node reboots.
Story 6: Emergency Hardware Removal (Self-Protection)
As a Node Operator, when my hypervisor unexpectedly reclaims physical memory from a running node (e.g., due to a balloon driver deflation or host pressure), I want the Kubelet to automatically detect the reduced capacity and immediately shrink its cgroup boundaries and eviction thresholds — protecting the node from kernel OOM panics — without requiring any manual intervention or Kubelet restart.
Story 7: Dynamic Storage Expansion
As a Storage Administrator, I want to dynamically expand the root block volume of a worker node on the fly, so that the Kubelet instantly recognizes the increased Ephemeral Storage capacity and allows pods to utilize the new space without triggering false disk-pressure evictions.
Note: Ephemeral storage resize follows the same declarative API model as CPU and Memory. The Kubelet’s capacity reconciliation loop updates
Node.Status.Capacity[ephemeral-storage]and the corresponding eviction thresholds whenNode.Spec.ConfiguredCapacityincludes an updated ephemeral storage value. Physical block-device expansion (e.g., resizing the underlying volume via a cloud provider) remains an external operation outside the scope of this KEP.
Story 8: Orchestrated Capacity Downscale
As a Cluster Administrator, I want to dynamically reclaim (hot-unplug) underutilized memory or CPU from a node without restarting the Kubelet. I want to orchestrate this via the Kubernetes API first, so workloads are gracefully evicted and the scheduler stops sending pods before I physically remove the hardware, avoiding node crashes and workload scheduling races.
Notes/Constraints/Caveats (Optional)
Linux and cgroup v2: This feature targets Linux nodes running cgroup v2. cgroup v1 nodes are not supported.
Swap-Enabled Nodes Not Supported (Alpha): In-place node resize is not supported on nodes with Swap enabled in the Alpha phase. If Swap is active on the node, the Kubelet will decline to perform a resize and will set the
CapacityConfiguredcondition toFalse(Reason:Infeasible) with a message indicating swap is unsupported. This restriction will be revisited in a subsequent milestone once the swap–pod-resize interaction is resolved.Linux Only: This feature has no effect on Windows nodes. The Kubelet’s capacity reconciliation loop short-circuits immediately on non-Linux platforms.
NUMA Topology Lazy Reconciliation: When a resize changes the available memory or CPU cores per NUMA zone, the Topology Manager’s view of NUMA boundaries is updated in its internal state machine. However, running pods retain their original NUMA pinning — they are not remapped mid-flight. Only newly admitted pods use the updated NUMA layout. Operators should account for this when sizing a downscale target on NUMA-pinned workloads.
Risks and Mitigations
OOMScoreAdjust Drift for Existing Pods
Risk: The Kubelet calculates a container’s
oom_score_adjupon creation using the formula:1000 - (1000 * containerMemoryRequest) / nodeMemoryCapacity. If a node’s memory capacity changes, the OOM scores of existing Burstable pods will mathematically drift compared to newly scheduled Burstable pods, potentially skewing the Linux OOM killer’s tie-breaker logic.Mitigation: We explicitly accept this minor drift. Updating
oom_score_adjfor running containers requires identifying and rewriting the/proc/<PID>/oom_score_adjfile for every single running thread inside the container. This introduces severe latency, high CPU overhead, and dangerous race conditions. The overarching QoS hierarchy (Guaranteed pods remain invincible, BestEffort pods remain first-to-die) is strictly preserved, making the risk of a slightly skewed Burstable tie-breaker acceptable compared to the danger of rewriting thousands of running PIDs.Container Swap Limit Re-calculation Overhead
Risk: The proportional swap limit for a container relies on the node’s total memory capacity. Upon a resize, failing to update this leads to stranded swap space (upscale) or immediate kernel panics (downscale). Recalculating and applying this to all active pods also introduces CRI overhead.
Mitigation: For the Alpha phase, this risk is eliminated entirely by not supporting resize on swap-enabled nodes (see Non-Goals). The Kubelet will decline to perform a resize if Swap is active, avoiding both the correctness and overhead concerns. The correct approach for swap recalculation — including aligning the semantics with pod resize (KEP-1287) — is deferred to a subsequent milestone.
Kubelet Sub-Manager Synchronization Failure
Risk: During an upscale, if the internal Kubelet sub-managers (CPU Manager, Memory Manager) fail to synchronize the new capacity, the Kubelet might reject new pod allocations, leading to underutilized hardware and scheduling deadlocks.
Mitigation: The reconciliation loop is self-healing and naturally guarded against Denial of Service (DoS) tight-looping. If a sub-manager fails to sync, the Kubelet aborts the cache update, emits an error, and increments the
kubelet_node_resize_errors_totalmetric. Because the internal capacity cache was not updated, the system will naturally re-attempt the reconciliation on the nextcAdvisorpolling cycle (defaulting to every 5 minutes) when the hardware drift is detected again. This strict 5-minute interval acts as a natural rate-limit, protecting the Kubelet’s CPU.API Status Clamping
Risk: The Kubelet’s node status updater relies on a boot-time MachineInfo cache. If the ContainerManager updates its internal Allocatable limits but fails to update this global cache, the status updater will aggressively clamp the new Allocatable value down to the stale boot-time Capacity, hiding the new hardware from the control plane forever.
Mitigation: The event-driven signal explicitly forces a refresh of kl.setCachedMachineInfo() before calling syncNodeStatus(). This ensures the status updater evaluates the new limits against the live physical reality, bypassing the clamping safeguard safely.
Application-Level Hardware Assumptions
Risk: Workloads often read /proc/cpuinfo or /proc/meminfo exactly once during their startup routine. If the node is vertically scaled, applications that spawn fixed per-CPU thread pools or rely on strict NUMA boundary alignments will not organically scale to use the new resources.
Mitigation: This is an accepted limitation and is treated as an application-level responsibility. Applications must be written to dynamically poll their limits (e.g., listening to cgroup file changes) or be manually restarted by their controlling Deployment to read the new hardware layout. The Kubelet’s responsibility is solely to make the hardware available at the cgroup boundary.
Coordination with External NRI/Runtime Plugins
Risk: External Node Resource Interface (NRI) plugins or custom runtime wrappers may cache node capacity independently of the Kubelet, leading to split-brain resource tracking after a resize event.
Mitigation: No new CRI or NRI notification call is introduced by this KEP. The runtime implicitly learns of the new capacity boundaries when the Kubelet pushes updated cgroup limits for each running container via the existing
UpdateContainerResourcesCRI RPC — the same mechanism used by In-Place Pod Resource Resize (KEP-1287). This means the container runtime and any NRI plugins that subscribe to cgroup changes will see the updated limits without a dedicated capacity-change event. Direct notification of NRI plugins via a new API is explicitly deferred to Future Work (see NRI Integration).Capacity Detection Latency (Downscale Risk)
Risk:
cAdvisorcaches and refreshesMachineInfoevery 5 minutes by default. During a memory hot-unplug (downscale) event, there is up to a 5-minute latency window where physical memory is removed, but the Kubelet remains unaware. If workloads spike during this window, the Linux kernel OOM killer may violently terminate processes before the Kubelet’s capacity reconciliation loop detects the hardware drop and triggers a gracefulNodeCapacityExceededeviction.Mitigation: For the Alpha phase, this detection latency is an accepted operational constraint, with the kernel OOM killer serving as the ultimate safety net to protect the node. For future phases (Beta/GA), the detection mechanism will transition to an event-driven architecture to eliminate this polling lag (see Future Work).
Hardware Metric Jitter and Malformed Signals
Risk: The Kubelet relies on
cAdvisorto read raw hardware metrics from the host operating system. The OS can occasionally report minor memory capacity jitter (micro-fluctuations in bytes) due to internal kernel allocations, or bugs could result in malformed data (e.g., negative capacities, or0). Blindly reacting to these would cause an endless loop of API patches, CRI overhead, and potential mass-evictions.Mitigation: Before accepting a capacity change, the ContainerManager implements a strict Validation and Jitter Tolerance Filter. Micro-drifts (e.g., changes under 100Mi) are silently dropped. Nonsense values (e.g., 0, negative values, or values that fall below the static Kubelet reservations) are explicitly rejected, an error is logged, and the reconciliation pipeline short-circuits.
Design Details
API Changes
A new declarative structure is introduced to the NodeSpec API, allowing external controllers or administrators to declare the desired logical capacity target for the node.
// 1. New field added to NodeSpec
type NodeSpec struct {
// ... existing fields ...
// ConfiguredCapacity defines the desired logical capacity of the node.
// On startup, the Kubelet checks this field first; if set, it is treated as
// the desired target and validated against the physical upper bound reported
// by cAdvisor. If unset, the Kubelet uses self-discovered cAdvisor capacity
// as both the physical upper bound and the initial target.
// For Alpha, ConfiguredCapacity must not exceed physical capacity (CPU/Memory).
// +optional
ConfiguredCapacity ResourceList `json:"configuredCapacity,omitempty"`
}
// 2. New Condition Type constant
const (
// ... existing conditions (e.g., NodeReady, NodeMemoryPressure) ...
// NodeCapacityConfigured indicates the status of the Kubelet's reconciliation
// of the desired Node.Spec.ConfiguredCapacity against the physical hardware.
NodeCapacityConfigured NodeConditionType = "CapacityConfigured"
)
Node Conditions for State Dissemination
To prevent Kubelet reconciliation loops and provide standard Kubernetes observability, the Kubelet exposes its validation and actuation state via a new Node Condition: CapacityConfigured. We intentionally mirror the state vocabulary established by In-Place Pod Resource Resize (KEP-1287).
Condition: CapacityConfigured
Status: True
Reason Accepted: The requested ConfiguredCapacity is valid, bounded by physical hardware constraints, and has been fully applied to local cgroups and Node.Status. The Kubelet considers the node stable and will not retry.
Status: False
Reason InProgress: The Kubelet has accepted a downscale request and is actively shrinking cgroups or gracefully evicting starved pods. The final Node.Status update is pending.
Reason Infeasible: The requested ConfiguredCapacity exceeds physical hardware limits (overcommit is disallowed in Alpha). The Kubelet has clamped the target to the physical limits, set this condition, and will not retry the oversized request. This provides explicit lockSize semantics: the Kubelet treats the Infeasible state as terminal for the current Spec value. An external controller or administrator must patch Node.Spec.ConfiguredCapacity to a valid, in-bounds value to resume normal reconciliation. A dedicated lockSize boolean field on NodeSpec was considered but rejected in favour of this condition reason — it provides the same terminal semantics without adding a new API field, and remains observable via standard kubectl get node condition output.
Reason EmergencyReduced: Physical hardware was forcefully removed (e.g., hypervisor-forced reclaim), falling below the current Node.Spec.ConfiguredCapacity. The Kubelet bypassed the API and clamped the node to the new physical reality to protect the kernel. Normal reconciliation is suspended until an external actor patches the Spec down to match the new physical bounds.
Admission Control Contract
Source of Truth:
The Node.Spec.ConfiguredCapacity field is the authoritative declaration of desired logical capacity. Node.Status.Capacity and Node.Status.Allocatable are read-only outputs of the Kubelet’s reconciliation — no external controller or webhook should patch them directly. The Kubelet is the sole component that writes to Node.Status.Capacity. This separation of Spec from Status follows the standard Kubernetes controller pattern.
Path A (Orchestrated): External controllers or administrators patch Node.Spec.ConfiguredCapacity. This mutation is intercepted by the cluster’s standard Validating and Mutating Webhooks. If a webhook rejects the resize request, the API Server denies the PATCH, and the Kubelet’s informer never receives the event — the node’s effective capacity does not change.
Path B (Emergency): When physical hardware is forcefully reclaimed (e.g., hypervisor-forced deflation), the Kubelet reacts to cAdvisor directly and patches Node.Status.Capacity and Node.Status.Conditions. This path utilizes the standard Node Authorizer RBAC, bypassing Spec webhooks to ensure the control plane is immediately notified of physical degradation without needing an external controller to be available.
Bootstrap Behavior: Upon restart, the Kubelet reads Node.Spec.ConfiguredCapacity from the API Server as the primary capacity target before reading the raw cAdvisor hardware data. If a valid ConfiguredCapacity exists in the Spec, the Kubelet treats it as the desired state and validates it against live physical hardware. This ensures that a Kubelet restart on a pre-configured node does not accidentally override the declared configuration.
Resource-Specific Validation Rules
Node.Spec.ConfiguredCapacity is a ResourceList — a map of resource name to quantity — and is entirely optional. Users are not required to set all resource types, or any at all:
- Field absent (
omitempty): The Kubelet behaves as today — it uses the raw cAdvisor-reported physical capacity for all resources and no reconciliation loop is started. - Field present, resource key absent: If
ConfiguredCapacityis set but does not include a particular resource (e.g., CPU is present but memory is not), the Kubelet treats the absent resource as having no declared target and continues to use the cAdvisor-reported physical capacity for that resource. Reconciliation for the absent resource is a no-op. - Field present, resource key set to zero: A zero value is treated as a malformed signal and rejected — the Kubelet sets
CapacityConfiguredtoFalse(Reason:Infeasible) and does not actuate the resize for any resource in that request.
ConfiguredCapacity is applied per-resource independently: setting a value for CPU does not implicitly affect the declared or effective capacity for memory, hugepages, or any other resource. Each key in the map is validated and reconciled in isolation.
The Alpha constraint (ConfiguredCapacity <= Physical Capacity) applies per-resource. The Kubelet validates the Spec against the host using the following resource-specific rules:
CPU & Memory: Strictly bounded by the physical hardware limits reported by cAdvisor. The ConfiguredCapacity for these resources must not exceed the raw physical quantity.
Hugepages: Validated against the pre-allocated hugepage pools configured at the OS kernel level (e.g., via /sys/kernel/mm/hugepages), not the total raw memory.
Ephemeral Storage: Validated against the available disk capacity as reported by the host OS. The ConfiguredCapacity for ephemeral-storage must not exceed the actual available disk space on the node’s root filesystem.
Unknown resource types: Any resource type present in ConfiguredCapacity that the Kubelet does not recognise (e.g., custom extended resources) is silently ignored by the validation loop. Only well-known resource types (CPU, Memory, Hugepages, Ephemeral Storage) are validated and actioned in Alpha.
Note on swap-enabled nodes: Swap itself is not a key in ConfiguredCapacity and is never set by users. The constraint is more subtle: when a node has Swap enabled and memory is resized, each running container’s per-container swap limit (memory.swap.max) must be recalculated — it is derived proportionally from the ratio of the container’s memory request to total node memory, multiplied by total available node swap. Changing node memory without updating these derived per-container limits leads to stranded swap space or kernel panics. In Alpha, memory resize on swap-enabled nodes is therefore not supported: if the node has Swap active and a memory value is present in ConfiguredCapacity, the Kubelet sets CapacityConfigured to False (Reason: Infeasible) and does not actuate the resize. The correct recalculation semantics are deferred to a subsequent milestone.
Security Considerations
This section explicitly addresses the security properties of this feature in response to the concern that a compromised node could over-report its capacity to the control plane in order to attract Pod scheduling and gain access to secrets it should not receive.
Capacity Inflation Attack is Prevented by Design (Alpha):
The Alpha enforcement rule (ConfiguredCapacity <= Physical Capacity) is the primary defense. The Kubelet’s calculateValidatedCapacity() function enforces this bound locally by reading raw capacity from cAdvisor, which reads directly from the host kernel (/sys, /proc). For a node to successfully inflate its reported capacity above its physical reality, an attacker would need to compromise either:
- The Kubelet binary itself, or
- The kernel-level data sources that
cAdvisorreads.
Both represent a full node compromise, which is already outside the Kubernetes threat model. A cluster-level actor patching Node.Spec.ConfiguredCapacity to an inflated value will have that spec clamped by the Kubelet and the CapacityConfigured condition set to False (Reason: Infeasible) — the API Server will store the spec, but the Kubelet will not act on it.
Node Authorizer and write access: Two distinct access-control mechanisms apply to Node.Spec.ConfiguredCapacity depending on the identity of the writer.
For Kubelet identities, the standard Node Authorizer enforces that each Kubelet may only write to the Node object that represents itself. This prevents a compromised node from patching ConfiguredCapacity on a neighbouring node — a Kubelet attempting to modify another node’s spec will be denied at the API Server by the Node Authorizer before the request reaches any webhook or admission plugin.
For non-Kubelet identities (external controllers, cluster administrators, automation), the Node Authorizer is not involved. Access is governed by standard RBAC: any identity with update or patch permission on the nodes resource can write ConfiguredCapacity. Cluster administrators who want to restrict which specific controllers are permitted to set this field can deploy a ValidatingWebhookConfiguration targeting Node.Spec.ConfiguredCapacity mutations.
NRI/Runtime Boundary:
The Kubelet does not introduce a new trust boundary between itself and the container runtime for capacity data. The runtime’s view of resource limits is updated via the existing UpdateContainerResources CRI call — the same path used by In-Place Pod Resource Resize (KEP-1287) — which carries no new elevation of privilege.
Architecture Flow
To safely support both external declarative triggers and physical hardware constraints, the Kubelet utilizes a dual-path reconciliation architecture based on the direction of the capacity change.
Path A1: Orchestrated Configuration-Driven Upscale
Spec mutation occurs before physical actuation, consistent with the API-first model.
1. Spec Mutation: The external controller patches Node.Spec.ConfiguredCapacity to the new higher target value. At this point, cAdvisor has not yet detected new hardware, so the Kubelet’s validation check (ConfiguredCapacity <= Physical Capacity) will temporarily fail. The Kubelet sets the CapacityConfigured condition to False (Reason: Infeasible) and holds the pending event in the reconciliation channel.
2. Physical Actuation: The external controller proceeds to physically increase capacity via the hypervisor hotplug. Once the hardware is online, cAdvisor detects the new physical capacity on its next poll cycle (default: 5 minutes, or triggered immediately by the cAdvisor hardware-change path).
3. Kubelet Reconciliation: The cAdvisor drift triggers the reconciliation goroutine. The validation check now passes: ConfiguredCapacity <= new Physical Capacity. The Kubelet clears the Infeasible hold.
4. Host Enforcement: The ContainerManager re-initializes sub-managers via the ResourceResizer interface, expands the host /kubepods cgroups, and updates CRI swap boundaries.
5. Status Dissemination: The Kubelet patches Node.Status.Capacity and transitions the CapacityConfigured condition to True (Reason: Accepted), advertising the new capacity to the Scheduler.
Path A2: Orchestrated Configuration-Driven Downscale
Logical API changes occur prior to physical hardware removal.
1. Spec Mutation: The external controller patches Node.Spec.ConfiguredCapacity to a lower value before making any physical alterations to the host VM.
2. Status Dissemination (Scheduler Block): The Kubelet detects the change. It immediately updates Node.Status.Capacity and Node.Status.Allocatable to reflect the downscale, physically preventing the Scheduler from assigning new pods to the node. It simultaneously sets the CapacityConfigured condition to False (Reason: InProgress), signaling to external controllers that the node is actively in a transient shrinking state.
3. Immediate Logic Enforcement: The ContainerManager instantly locks internal state, shrinks the /kubepods host cgroups, and recalculates absolute Eviction Manager thresholds against the new logical baseline.
4. Graceful Workload Degradation: The Kubelet evaluates active workloads against the newly reduced capacity. If strict request contracts can no longer be met, starved pods are gracefully evicted with Reason: NodeCapacityExceeded.
5. State Transition (Accepted): Once evictions are complete and the node’s boundaries are fully secured, the Kubelet transitions the CapacityConfigured condition to True (Reason: Accepted) and pushes the final status update.
6. Physical Reclaim: The external controller observes the Accepted condition and safely executes the physical hardware hot-unplug via the hypervisor.
Path B: Emergency Hardware-Driven Fallback (Ungraceful Downscale)
When physical hardware (e.g., Memory) is forcefully yanked by a hypervisor without prior API synchronization, physics dictates the timeline. The Kubelet acts as an emergency circuit breaker to secure the host.
1. Hardware Trigger: cAdvisor detects a drop in physical metrics that falls below the current Node.Spec.ConfiguredCapacity value.
2. Immediate Host Enforcement: The ContainerManager bypasses the API state and instantly shrinks the /kubepods cgroups to match the raw physical limits.
3. Emergency Eviction: The Kubelet evaluates active workloads against the raw physical bounds and immediately evicts starved pods.
4. State Override & Alerting: The Kubelet calculates the clamped target and patches Node.Status.Capacity. It transitions the CapacityConfigured Condition to False (Reason: EmergencyReduced) and emits a Warning Event (EmergencyCapacityReduced). To prevent infinite API loops, the Kubelet caches this clamped target locally and safely ignores the oversized Spec until the Spec is updated by an external actor to match or fall below the new physical reality.
5. Divergence Resolution: The external controller watches for the EmergencyReduced condition. Upon seeing it, the controller is responsible for patching Node.Spec.ConfiguredCapacity down to match reality, which clears the split-brain state and resumes normal Kubelet reconciliation behavior.
Flow Control: Container Swap Limit Recalculation
Alpha scope note: Resize on swap-enabled nodes is not supported in the Alpha phase (see Non-Goals). This section describes the intended design for a future milestone when swap support is introduced.
If a node is configured with Swap, a container’s swap limit is dynamically proportional to the total node memory. Failing to update this during a resize leads to stranded resources or immediate kernel panics.
Formula: (<containerMemoryRequest> / <nodeTotalMemory>) * <totalPodsSwapAvailable>
T=0: Initial Node Resources
- Node Memory: 6G
- Node Swap: 4G
Pod (Running):
- MemoryRequest: 2G
Runtime Cgroup (<cgroup_path>/memory.swap.max): 1.33G
T=1: Hot-plug Memory (Upscale)
- Node Memory: 8G (+2G)
- Node Swap: 4G
Pod (Running):
- MemoryRequest: 2G
Runtime Cgroup (<cgroup_path>/memory.swap.max): 1.0G (Recalculated via CRI)
Flow Control: Hardware Degradation and Capacity Starvation
During a hot-unplug (downscale) event, the node’s physical capacity may drop below the total resources currently requested by active workloads. Because the Kubelet can no longer fulfill the strict scheduling contract, it must intervene to prevent system lockups or kernel panics.
- Memory Starvation: If the newly reduced node memory capacity falls below the aggregate memory requests or active working set of running pods, the Eviction Manager will trigger standard memory-pressure evictions.
- CPU Starvation (Guaranteed QoS): If the node’s CPU core count drops below the threshold required to fulfill the exclusive core allocations of
Guaranteedpods (managed by thestaticCPU Manager policy), the Kubelet cannot safely throttle the workload without violating the SLA.
In these starvation scenarios, the Kubelet’s eviction manager will gracefully terminate the affected pods with a Failed status (Reason: NodeCapacityExceeded). This explicitly forces the cluster-wide controllers (e.g., Deployments, StatefulSets) to immediately reschedule the workload onto a capable, healthy node.
Note on Static Pods: Static pods are managed directly by the Kubelet via local manifest files and are not subject to eviction manager decisions. They will not be evicted during a capacity downscale. Cluster operators must manually account for the resource footprint of any static pods when setting a downscale target in Node.Spec.ConfiguredCapacity, ensuring the declared target leaves sufficient headroom above the aggregate static pod requests.
Compatibility with Cluster Autoscaler
The Cluster Autoscaler (CA) presently anticipates uniform allocatable values among nodes within the same NodeGroup, using existing nodes as templates for newly provisioned nodes. With mutable node capacity, nodes within a single group may drift in size over time, which can cause the CA to select a resized node as its provisioning template — causing it to expect new nodes to have the larger capacity, while the cloud provider provisions base-sized nodes, leading to scheduling failures.
This is an open design problem. Placing a static boot-time value on the Node object (as an annotation or a dedicated status field) has limitations: Node.Status is intended to reflect live state rather than historical boot state; annotations are brittle when multiple controllers read and react to them; and if every node in a NodeGroup has been uniformly resized, the “initial” capacity is no longer a meaningful reference — the NodeGroup has functionally changed size and operators may reasonably expect that to be reflected.
Alternative approaches include having CA read directly from the cloud provider’s launch template (which reflects the true provisioning baseline) or introducing a configurable reference capacity within CA’s own NodeGroup configuration. Both decouple the problem from the Kubernetes Node object and place it where it belongs — in the component that understands provisioning semantics.
The Layer 1 deliverable for CA compatibility is to document the current behaviour and the failure mode, and validate that CA correctly observes Node.Status.Capacity UPDATE events when capacity changes. The right long-term fix will be agreed upon before any CA integration is standardised in this KEP.
Layer 1 Implementation: Ecosystem Tolerance — Validating That Mutable Capacity Is Safe
This section implements the Layer 1 contracts described in the Proposal. It formally documents the static-capacity assumptions that Kubernetes currently holds, and adds the upstream test coverage that proves the ecosystem can safely handle a Node.Status.Capacity mutation — whether that mutation arrives via a Kubelet restart or via the Layer 2 reconciliation loop.
Assumptions Being Changed
- Static Capacity Assumption:
Node.Status.CapacityandAllocatableare currently fixed for the lifetime of a running Kubelet process — they are written at boot and held constant until the next restart. This KEP transitions them to fields that can be mutated on a live,Readynode via the declarative reconciliation loop. That live-mutation path has never been formally tested or supported by the ecosystem. - Kubelet Restart Admission: Currently, if a Kubelet is restarted on a machine whose hardware was reduced while offline, the Kubelet may blindly fail pod admission. We are shifting this to a graceful reconciliation and eviction model, whether the trigger is a restart or the live reconciliation loop.
- Autoscaler Homogeneity: Nodes within a single NodeGroup will no longer be guaranteed to have identical capacities, meaning the Autoscaler cannot blindly select any node as a provisioning template.
- External Controller Caches: Third-party operators that cache node sizes indefinitely will become stale. (This is an accepted operational constraint).
Layer 1 Pre-requisite Tests
These tests prove the foundational ecosystem contracts hold. They are deliberately scoped to raw API and control-plane behaviour — no Kubelet reconciliation code is exercised. Passing these tests is the gate for beginning Layer 2 implementation.
Test 1: API Server Mutation Acceptance
- Action: Manually patch
.status.capacityand.status.allocatableon aReadyNode object. - Validation: Verify the API Server accepts the patch without systemic webhook rejections or validation failures.
- Layer: Layer 1 — API Server contract.
- Action: Manually patch
Test 2: Scheduler Cache Invalidation (Upscale)
- Action: Create a pending Pod that requires 8Gi of memory on a cluster where the only node has 4Gi. Manually patch the Node object’s capacity to 10Gi.
- Validation: Verify the Scheduler detects the mutated Node object, updates its internal cache, and successfully schedules the pending Pod.
- Layer: Layer 1 — Scheduler contract.
Test 3: Scheduler Cache Invalidation (Downscale)
- Action: Manually patch an empty Node’s capacity from 10Gi down to 4Gi. Attempt to schedule a Pod requiring 8Gi.
- Validation: Verify the Scheduler respects the mutated smaller capacity and rejects the Pod (leaves it Pending), proving it does not rely on a stale boot-time cache.
- Layer: Layer 1 — Scheduler contract.
Test 4: Kubelet Restart on Resized Hardware (Workaround Boundary)
- Action: Schedule a Pod. Stop the Kubelet. Mock the underlying machine info to reflect a smaller capacity (simulate offline hot-unplug). Start the Kubelet.
- Validation: Verify the Kubelet boots successfully, recognizes the discrepancy between the API and physical hardware, and handles the change gracefully (e.g., evicting the pod if starved) rather than crashing or permanently locking pod admission.
- Layer: Layer 1 — establishes the Kubelet-restart workaround boundary: the minimum safe behaviour that Layer 2 must meet or exceed.
Test 5: Deferred Pod Resize Retried on Node Upscale
- Action: On a node with 4Gi allocatable memory, submit a pod with a pending resize request to 3.5Gi (which is
Deferredbecause the node is fully packed by other pods). Manually patchNode.Status.Allocatableupward to 8Gi, freeing headroom. - Validation: Verify the Kubelet re-evaluates the previously
Deferredresize and transitions it toInProgress/Acceptednow that sufficient allocatable capacity exists. Verify the pod’sresizestatus condition reflects the updated state. - Layer: Layer 1 — KEP-1287 interaction contract (upscale path).
- Action: On a node with 4Gi allocatable memory, submit a pod with a pending resize request to 3.5Gi (which is
Test 6: Deferred Pod Resize Transitions to Infeasible on Node Downscale
- Action: On a node with 8Gi allocatable memory, submit a pod with a pending resize request to 6Gi (marked
Deferreddue to contention). Manually patchNode.Status.Allocatabledownward to 3Gi — below the desired resize target. - Validation: Verify the Kubelet detects that the deferred resize can no longer fit and transitions the pod’s resize status from
DeferredtoInfeasible. Verify the Scheduler and VPA observe the updated status (they should no longer treat this as a pending retry). - Layer: Layer 1 — KEP-1287 interaction contract (downscale path).
- Action: On a node with 8Gi allocatable memory, submit a pod with a pending resize request to 6Gi (marked
Test 7: Infeasible Pod Resize Cleared on Node Upscale
- Action: On a node with 4Gi allocatable memory, submit a pod resize request to 6Gi. The Kubelet marks it
Infeasibledue to insufficient node capacity. Manually patchNode.Status.Allocatableupward to 8Gi. - Validation: Verify the Kubelet re-evaluates the
Infeasibleresize, determines the node now has sufficient capacity, and transitions the resize status toInProgress/Accepted. Confirm the Kubelet correctly distinguishes between a capacity-boundInfeasible(retriable) and anInfeasiblecaused by other reasons (not retriable, e.g., resource type not supported). - Layer: Layer 1 — KEP-1287 interaction contract (capacity-bound Infeasible retry).
- Action: On a node with 4Gi allocatable memory, submit a pod resize request to 6Gi. The Kubelet marks it
Test 8: Scheduler Preemption Recalculation on Capacity Decrease During Grace Period
- Action: On a node with 8Gi allocatable memory running a low-priority pod consuming 6Gi, submit a high-priority pod requiring 7Gi. The Scheduler selects the low-priority pod as a preemption victim and initiates its termination grace period. Before the grace period expires, manually patch
Node.Status.Allocatabledownward to 4Gi. - Validation: Verify the Scheduler receives the Node UPDATE event, re-queues the waiting high-priority pod for a fresh scheduling cycle, and re-runs preemption calculations against the new (4Gi) allocatable value — rather than proceeding with the stale assumption that 8Gi will be available once the victim terminates. Confirm the queueing hint for
Node.Status.Allocatabledecreases correctly triggers re-evaluation of pods in theWaitingForPreemptionstate. - Layer: Layer 1 — Scheduler preemption queueing hint contract.
- Action: On a node with 8Gi allocatable memory running a low-priority pod consuming 6Gi, submit a high-priority pod requiring 7Gi. The Scheduler selects the low-priority pod as a preemption victim and initiates its termination grace period. Before the grace period expires, manually patch
Layer 2 Implementation: Declarative Capacity Actuation
This section implements the Layer 2 changes described in the Proposal: the Node.Spec.ConfiguredCapacity API field, the Kubelet reconciliation loop, and the CapacityConfigured condition. These changes are gated on Layer 1 pre-requisite tests passing, since the reconciliation loop produces the same kind of Node.Status.Capacity mutation that Layer 1 validates the ecosystem can handle safely.
Proposed Core Code Changes
Dedicated Node Capacity syncLoop (
pkg/kubelet/kubelet.go)A dedicated node capacity syncLoop runs as its own goroutine within
Run(), analogous to the pod syncLoop. It owns the full lifecycle of a capacity reconciliation event and executes all reconciliation steps serially, ensuring no concurrent capacity mutations. The loop is fed by a singlecapacityReconciliationChchannel written by two sources: the Node informer (whenNode.Spec.ConfiguredCapacitychanges) and the cAdvisor hardware-drift detector (when physical capacity changes).
// 1. Wire up the Node Informer to feed the capacity syncLoop
kl.nodeInformer.AddEventHandler(cache.ResourceEventHandlerFuncs{
UpdateFunc: func(oldObj, newObj interface{}) {
oldNode := oldObj.(*v1.Node)
newNode := newObj.(*v1.Node)
if !apiequality.Semantic.DeepEqual(oldNode.Spec.ConfiguredCapacity, newNode.Spec.ConfiguredCapacity) {
kl.capacityReconciliationCh <- struct{}{}
}
},
})
// 2. Node capacity syncLoop
if utilfeature.DefaultFeatureGate.Enabled(features.InPlaceNodeResourceResize) {
go wait.Until(func(ctx context.Context) {
// capacityReconciliationCh is written by: Node informer OR cAdvisor hardware-drift detector
for range kl.capacityReconciliationCh {
configuredCapacity := kl.getNodeSpecConfiguredCapacity()
physicalInfo, err := kl.cadvisor.MachineInfo()
if err != nil { continue }
// Determine Target Capacity (Alpha rule: Configured <= Physical)
targetCapacity, condition := kl.calculateValidatedCapacity(configuredCapacity, physicalInfo)
// Anti-Thrash Guard: Break infinite loops if Spec > Physical
if kl.requiresReconciliation(targetCapacity) {
// Refresh internal cache to the validated target capacity
kl.setCachedMachineInfo(targetCapacity)
// CRITICAL DOWNSCALE FIX: Update Status and Allocatable FIRST
// to prevent Scheduler livelocks during the eviction window.
kl.updateNodeCondition(InProgress)
kl.syncNodeStatus(ctx)
// Enforce Host Boundaries (cgroups, sub-managers)
kl.containerManager.SyncCapacity(targetCapacity)
// Sync Eviction Thresholds for safe downscaling
kl.evictionManager.SynchronizeThresholds(targetCapacity)
// Gracefully evict pods starved by the new bounds
kl.evictStarvedPods(ctx, targetCapacity)
// Update active container Swap boundaries via standard CRI RPC
for _, pod := range kl.GetActivePods() {
kl.syncContainerSwapLimits(ctx, pod, targetCapacity)
}
// Disseminate final resolved state (Accepted, Infeasible, or EmergencyReduced)
kl.updateNodeCondition(condition)
kl.syncNodeStatus(ctx)
}
}
}, 0, wait.NeverStop)
}
The Sub-Manager Interfaces (
pkg/kubelet/kubelet.go)Resource managers must implement a new interface to accept dynamic sync events natively on the running node.
// ResourceResizer defines the interface for sub-managers to accept dynamic capacity changes
type ResourceResizer interface {
// SyncCapacity safely re-evaluates the sub-manager's internal state against the new boundaries
SyncCapacity(capacity ResourceList) error
}
- Resource Manager Synchronization and State Reconciliation
When the Container Manager detects a capacity drift, it notifies its sub-managers to synchronize their internal state. Because the Kubelet manages cpusets, memory pages, checkpointed state files, and NUMA alignments, dynamic resize requires coordinated reconciliation across each subsystem:
3.1 CPU Manager Synchronization
The CPU Manager dynamically reconciles capacity changes depending on the configured policy:
Policy:
none- All pods run across the entire machine’s cpuset. On upscale or downscale, the Kubelet updates the host and QoS cgroup cpuset hierarchies to match the new root cpuset.
Policy:
static- Shared Pool Updates for Burstable and BestEffort Pods:
- Non-Guaranteed pods (Burstable, BestEffort) and Guaranteed pods with non-integer CPU requests execute within the default shared cpuset pool (all physical CPUs excluding reserved CPUs and active exclusive allocations).
- When capacity changes, the CPU Manager recalculates the shared cpuset pool and triggers an active reconciliation run to push updated cpuset boundaries to all running Burstable and BestEffort containers through standard container runtime resource updates.
- Reserved CPU Invariant:
- Reserved CPUs represent an invariant reservation for host and kubelet system daemons.
- If a hot-unplug event attempts to remove CPU IDs that overlap with the configured reserved CPUs, the Kubelet rejects the downscale as infeasible to protect host system stability.
- Guaranteed Pods and Exclusive Core Allocations:
- Upscale: Newly added CPU IDs expand the shared pool, making more cores available for shared workloads or for subsequent Guaranteed pod admissions.
- Downscale: If CPU core removal reduces total capacity below the count required for active exclusive allocations, or if removed CPU IDs directly overlap with exclusive cores pinned to running Guaranteed containers, the Kubelet evicts the affected pods with a
Failedstatus (Reason:NodeCapacityExceeded).
- Burstable Pod Degradation Semantics:
- For Burstable pods, CPU requests establish scheduler bandwidth shares. CPU is a compressible resource: if a downscale reduces node CPU below aggregate Burstable requests, the Kubelet’s eviction manager does not proactively evict those pods — they gracefully degrade and share available CPU bandwidth proportionally across the contracted shared cpuset.
- However, this does not prevent Scheduler-driven preemption. When new higher-priority pods need to be scheduled onto the node after a downscale, the Scheduler sums aggregate requests against the new (smaller) allocatable value and will preempt lower-priority Burstable pods to make room, following standard Kubernetes preemption semantics. This KEP introduces no changes to that behaviour: preemption decisions remain entirely within the Scheduler and are driven by PriorityClass, not by this reconciliation loop.
- Shared Pool Updates for Burstable and BestEffort Pods:
3.2 Memory Manager and Memory QoS Synchronization
- NUMA Node Allocation Boundaries:
- The Memory Manager re-evaluates available physical memory and hugepages per NUMA node, updating its internal state memory map.
- Memory QoS:
- For nodes running with Memory QoS enabled, resizing node allocatable memory alters the proportional calculation for memory protection and throttling boundaries.
- If a memory downscale causes aggregate memory usage to exceed new limits, standard Memory QoS throttling triggers, followed by standard Eviction Manager ranking (evicting BestEffort workloads before Burstable).
3.3 Topology Manager and NUMA Layout
- Machine Topology Refresh:
- The Topology Manager queries the refreshed machine topology information to update its internal NUMA cell map, socket counts, and distance matrix.
- Admission Alignment:
- Running pods retain their existing NUMA node and resource pinning. Future pod admissions use the updated NUMA boundaries and refreshed hint providers for single-NUMA or multi-NUMA alignment decisions.
3.4 Checkpoint State File Consistency
- The CPU Manager and Memory Manager persist state across restarts via local checkpoint state files on disk.
- Capacity synchronization ensures that whenever in-memory state (such as the shared pool, allocations, and NUMA memory maps) is modified, the new topology and allocation table are atomically committed to the state files on disk. This prevents topology validation errors during subsequent Kubelet restarts.
3.5 Feature Scope Progression (Alpha to GA)
To ensure safety and manage implementation complexity:
- Alpha Scope:
- CPU Manager: Full support for
cpuManagerPolicy: none. ForcpuManagerPolicy: static, support upscaling (expanding the shared cpuset pool) and non-destructive downscaling (reclaiming unallocated shared cores). Hot-unplugging cores that conflict with reserved CPUs or allocated exclusive cpusets is rejected. - Memory Manager: Scoped to
memoryManagerPolicy: None. - Topology Manager: Scoped to
topologyManagerPolicy: noneor single-NUMA architectures. - Swap: Not supported. Resize operations on swap-enabled nodes are rejected in Alpha.
- CPU Manager: Full support for
- Beta Scope:
- Swap-enabled node resize, including per-container swap limit recalculation via
UpdateContainerResourcesCRI RPC, once swap–pod-resize interaction semantics are aligned. - Dynamic multi-NUMA topology cell changes, multi-NUMA memory block redistribution (
memoryManagerPolicy: Static), and full Topology Manager hint provider recalculation across dynamic NUMA boundaries.
- Swap-enabled node resize, including per-container swap limit recalculation via
Observability and Metrics
To ensure cluster operators can monitor resize events and failures, this KEP introduces the following Prometheus metrics within the Kubelet:
kubelet_node_resize_requests_total: Counter tracking the number of successful native resource resize events (labeled by resource name and direction:increase/decrease).kubelet_node_resize_errors_total: Counter tracking failures during the reconciliation pipeline (labeled by the failing subsystem, e.g.,cpu_manager_sync,cgroup_update).
Test Plan
[x] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.
Layer 1: Ecosystem Tolerance Tests (Pre-requisite, no feature gate required)
These tests must be merged before any Layer 2 Kubelet code is written. They verify that the control plane safely handles a Node.Status.Capacity mutation regardless of how it was triggered:
- API Validation: Verify that modifying
Node.Status.CapacityandNode.Status.Allocatableon an existing Node object is explicitly permitted by the API server and does not trigger unintended systemic webhook rejections. - Scheduler Validation: Verify that if a Node’s capacity is mutated (simulating an offline/restarted Kubelet resize), the Scheduler correctly recognizes the new capacity and successfully schedules/rejects pending Pods accordingly without requiring the Node object to be deleted and recreated.
- Autoscaler Integration: Verify how the Cluster Autoscaler reacts to a dynamically mutated
Node.Status.Capacityand document the current behaviour — specifically whether a resized node gets incorrectly selected as a NodeGroup provisioning template. The correct long-term fix for CA template stability is an open design item (see Cluster Autoscaler compatibility discussion above).
Layer 2: Unit tests (require InPlaceNodeResourceResize feature gate)
cAdvisor Cache Refresh (
kubelet_node_status_test.go): Verify that injecting a newMachineInfostruct successfully updates the Kubelet’s internal cache, and that the subsequent node status sync does not incorrectly clamp the newAllocatablevalues to the old boot-time capacity.Cgroup Enforcement (
container_manager_linux_test.go): Verify that when a capacity change is detected,enforceNodeAllocatableCgroupsandUpdateQOSCgroupsare invoked with the newly calculated boundaries, and that thenodeCapacityUpdateChsignal is successfully emitted without blocking.Sub-Manager Re-initialization (
cpu_manager_test.go,memory_manager_test.go): Verify that the CPU and Memory managers properly implement theResourceResizerinterface and cleanly acceptSyncCapacity()calls without leaking state or crashing.Eviction Threshold Sync (
eviction_manager_test.go): Verify thatSynchronizeThresholdscorrectly recalculates absolute byte values (e.g., <100Mivs10%) when the underlying capacity increases or decreases.CRI Swap Limit Recalculation (
kubelet_test.go): Verify the math for proportional swap limits. Ensure the capacity reconciliation loop correctly iterates over active pods and invokes the mock CRIUpdateContainerResourcesinterface with the newly calculated boundaries.Bootstrap Parity (
kubelet_test.go): Verify that when the Kubelet starts and finds a pre-existingNode.Spec.ConfiguredCapacityin the API Server, it uses that value as the desired target rather than silently overwriting it with the raw cAdvisor-discovered capacity. Specifically confirm that ifConfiguredCapacityis set to 20Gi on a 32Gi physical node, the Kubelet reports 20Gi inNode.Status.Capacityafter startup, not 32Gi.
Layer 2: e2e tests (require InPlaceNodeResourceResize feature gate)
These tests will utilize a mock cAdvisor interface to inject dynamic hardware capacity changes into a running test Kubelet to validate the end-to-end reconciliation pipeline.
Scenario 1: Safe Upscale and Scheduling
Action: Inject an upscale event (e.g.,
10G->15Gmemory).Validation: Verify the
/kubepodshost cgroup expands. Verify the Node API object reflects the new capacity. Verify a previouslyPendingpod (due to lack of memory) is successfully scheduled and transitions toRunning.
Scenario 2: Safe Downscale and Eviction (Resource Starvation)
Action: Schedule pods that consume 8G of memory. Inject a downscale event reducing the node’s total memory to 5G.
Validation: Verify the Kubelet updates its Eviction Manager thresholds and successfully evicts the lowest-priority pod (e.g.,
BestEffort) to protect the node before the cgroups enforce the 5G limit.
Scenario 3: Proportional Swap Recalculation
Action: Deploy a pod on a swap-enabled node. Inject a memory upscale event.
Validation: Inspect the active pod’s
memory.swap.maxcgroup file on the host filesystem and verify the limit was proportionally reduced based on the new total node memory.
Scenario 4: Upsize -> Downsize -> Upsize (Flapping)
Action: Rapidly inject alternating capacity changes.
Validation: Ensure the
ContainerManagerreconciliation loop does not deadlock, the cgroups settle on the final capacity, and no duplicate capacity update signals block the main Kubelet loop.
Graduation Criteria
Phase 1: Alpha (v1.38) — Layer 1: Ecosystem Tolerance
The v1.38 Alpha milestone is scoped to Layer 1 only. The goal is to prove the broader control-plane ecosystem can safely tolerate a live change to Node.Status.Capacity — using the Kubelet-restart-on-resized-hardware path as the initial trigger — before any new declarative API or live Kubelet actuation (Layer 2) is introduced.
- Layer 1 (Ecosystem Tolerance): API, Scheduler, and Autoscaler e2e tests are merged to officially validate and document that the Kubernetes ecosystem can safely handle
Node.Status.Capacitymutations. These tests have no feature gate dependency and establish the safety baseline for all Layer 2 work. Specifically:- The API Server accepts live patches to
Node.Status.Capacityon aReadynode. - The Scheduler’s
NodeInfocache correctly invalidates and re-evaluates capacity on Node UPDATE events (both upscale and downscale). - The Cluster Autoscaler’s NodeGroup template behaviour when
Node.Status.Capacityis mutable is documented and the failure mode is validated. The long-term CA fix is an open design item. - VPA recommendations are re-bounded against updated node allocatable values.
- ResourceQuota enforcement is not bypassed by a capacity change.
- The API Server accepts live patches to
- Kubelet Restart Boundary: The Kubelet correctly handles a restart on a node whose hardware capacity changed while offline — reconciling gracefully rather than crashing or permanently blocking pod admission (Test 4 in the Layer 1 pre-requisite test plan).
- No new
NodeSpecAPI fields are introduced in this milestone. NoInPlaceNodeResourceResizefeature gate is required for any v1.38 deliverable.
Phase 2: Alpha (v1.39) — Layer 2: Declarative Capacity Actuation (planned)
Layer 2 targets a subsequent milestone once the Layer 1 ecosystem contracts are validated and the open design questions in Open Questions for Layer 2 are resolved. The criteria below are provisional.
- Feature is disabled by default via the
InPlaceNodeResourceResizefeature gate (kubelet, kube-apiserver, kube-scheduler). - The
Node.Spec.ConfiguredCapacityAPI field is introduced. The Kubelet reconciles capacity mismatches for CPU and Memory — covering both the live-node and Kubelet-restart-on-resized-hardware cases. The Kubelet updates cgroups, re-initialises sub-managers (cpuManagerPolicy: none,memoryManagerPolicy: None), and evicts starved pods. - The cAdvisor metrics-based hardware-drift trigger is implemented, enabling the emergency downscale path (Path B).
- The Scheduler Node UPDATE queueing hint for
WaitingForPreemptionpods is implemented and validated. - All Layer 1 pre-requisite tests continue to pass.
- Integrations with
cpuManagerPolicy: static, Swap, and Topology Manager are explicitly deferred to Beta. - Out-of-tree controller integrations (Cluster Autoscaler, VPA) are not required for this milestone.
Phase 3: Beta
- The
InPlaceNodeResourceResizefeature gate is enabled by default. - Memory resize on swap-enabled nodes is supported: per-container
memory.swap.maxrecalculation viaUpdateContainerResourcesCRI RPC is implemented and validated. cpuManagerPolicy: staticresize is fully supported for both upscale (expanding the shared cpuset pool) and non-destructive downscale; destructive downscale evicts affected pods withNodeCapacityExceeded.- Topology Manager integration is complete for multi-NUMA configurations: topology cell changes,
memoryManagerPolicy: Staticredistribution, and hint provider recalculation across dynamic NUMA boundaries. - The Cluster Autoscaler NodeGroup template stability problem is resolved (see Open Questions for Layer 2 ) and the agreed approach is implemented.
- The Scheduler
min(Node.Spec.ConfiguredCapacity, Node.Status.Capacity)capacity view, or the equivalent mitigation agreed in Open Question 3, is implemented and validated. - Rollout, upgrade, and rollback planning is completed (required for Beta PRR).
Phase 4: GA
- The
InPlaceNodeResourceResizefeature gate is removed (always on). - The feature has been enabled by default for at least two release cycles with no regressions.
- All e2e tests are flake-free for a minimum two-week window and meet Conformance Test requirements.
- All open design questions from the Alpha/Beta period are resolved and documented.
Upgrade / Downgrade Strategy
Upgrade
For the v1.38 Alpha (Layer 1), no Kubelet restart or feature gate change is required. The Layer 1 tests operate purely at the control-plane level and are always enabled.
For the v1.39 Alpha (Layer 2), the Kubelet must be restarted with the InPlaceNodeResourceResize feature gate enabled. Existing clusters do not experience any immediate impact upon upgrade; the Kubelet will simply begin mirroring the existing physical hardware capacity dynamically.
Downgrade
For Layer 2, it is trivially possible to downgrade by disabling the InPlaceNodeResourceResize feature gate and restarting the Kubelet. The Kubelet will revert to its legacy behavior: capturing the node capacity once during boot and freezing it. Any subsequent hardware hot-plugs will be safely ignored. The Layer 1 ecosystem tests are always-on and have no rollback requirement.
Version Skew Strategy
The interaction between the Kubelet and the control plane (specifically the Scheduler) relies entirely on standard Node API update events. When the Kubelet patches the Node.Status.Capacity and Allocatable fields, the scheduler’s existing node update event handler seamlessly processes the altered capacity.
Because this leverages pre-existing API contracts, no special coordination or version skew mitigation is required between the Kubelet and the control plane. Similarly, no updates are required for CRI, CNI, or CSI plugins prior to enabling this Kubelet feature.
Scheduler-Kubelet Race Window: A capacity change, like a node going NotReady, can invalidate an in-flight scheduling decision. This race window is inherent to the Kubernetes scheduler’s optimistic concurrency model and is not unique to this KEP. Specifically, for systems using Workload Aware Scheduling (WAS), a capacity downscale occurring between the scheduler’s binding decision and the Kubelet’s admission check may cause a pod rejection. The Kubelet will return an admission failure, and the pod will be rescheduled by its controlling workload controller. This behavior is consistent with existing failure handling in the scheduler and does not require changes to the scheduler for Alpha.
To minimise this window during a downscale, the Scheduler uses min(Node.Spec.ConfiguredCapacity, Node.Status.Capacity) as its view of effective node capacity. As soon as an operator writes a reduced value to Node.Spec.ConfiguredCapacity, the Scheduler conservatively accounts for the smaller capacity in its scheduling decisions — even before the Kubelet has finished reconciling Node.Status.Capacity. This is directly analogous to the treatment for in-place pod resize, where the Scheduler uses max(pod.spec.resources, pod.status.resources) to avoid scheduling against a pod whose resources are still being expanded. Together, these two conventions encode a consistent principle: always assume the worst-case resource footprint for any in-flight change. The race window is reduced to the interval between the Node.Spec.ConfiguredCapacity write and the Scheduler’s next cache refresh, rather than the full duration of Kubelet reconciliation. Kubelet admission remains the authoritative gate and the final safety net for correctness.
Production Readiness Review Questionnaire
Feature Enablement and Rollback
How can this feature be enabled / disabled in a live cluster?
- Feature gate (also fill in values in
kep.yaml)- Feature gate name:
InPlaceNodeResourceResize - Components depending on the feature gate:
kubelet,kube-apiserver,kube-scheduler(Layer 2 only; no feature gate is required for the v1.38 Alpha Layer 1 work) - Will enabling / disabling the feature require downtime of the control plane? For the v1.38 Alpha (Layer 1), no feature gate exists and no component restart is required — the Layer 1 work consists entirely of test coverage and has no runtime impact. For the v1.39 Alpha (Layer 2), enabling
InPlaceNodeResourceResizerequires a coordinated rollout across all three gated components (kubelet, kube-apiserver, kube-scheduler). The API Server must be updated first so the newNode.Spec.ConfiguredCapacityfield is recognised before the Kubelet begins writing to it. The Scheduler must be updated to activate the Node UPDATE queueing hint and themin()capacity view. Rolling the control plane during the upgrade constitutes the required downtime; it is bounded to the standard control-plane rolling-update window and does not affect running workloads. - Will enabling / disabling the feature require downtime or reprovisioning of a node? Yes, a Kubelet restart is required to toggle the
InPlaceNodeResourceResizefeature gate on the node. This does not disrupt running pods; existing cgroups and container state are preserved by the container runtime across the restart.
- Feature gate name:
Does enabling the feature change any default behavior?
No immediate behavior changes occur if the underlying node hardware has not changed. If the hardware does change, the default behavior changes from “ignoring the hardware change” to:
Upscale: Dynamically patching the
NodeAllocatable resources and rewriting host cgroups, allowing pending pods to be scheduled.Downscale: Dynamically shrinking host cgroups, lowering Eviction Manager thresholds, and potentially triggering pod evictions.
Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?
Yes. The feature can be disabled by restarting the Kubelet with the feature gate turned off. The Kubelet will freeze its capacity at whatever cAdvisor reported during that specific boot cycle.
What happens if we re-enable the feature if it was previously rolled back?
The Kubelet will immediately poll the live cAdvisor data, detect any drift that occurred while the feature was disabled, and execute a one-time reconciliation to update the internal cgroups and the API Server’s Node.Status.
Are there any tests for feature enablement/disablement?
Yes, unit tests will validate that the capacityReconciler Go routine completely short-circuits and exits if utilfeature.DefaultFeatureGate.Enabled() returns false.
Rollout, Upgrade and Rollback Planning
How can a rollout or rollback fail? Can it impact already running workloads?
Rollout failures are isolated to the specific node. If the ContainerManager fails to enforce the new top-level /kubepods cgroups due to a filesystem error, the main Kubelet loop will not be signaled, and the API Server will not be updated. Existing running workloads are perfectly safe and will continue to operate under their current cgroup limits.
What specific metrics should inform a rollback?
An operator should roll back the feature if there is a sustained spike in the kubelet_node_resize_errors_total metric, indicating the Kubelet is deadlocking or failing to write to the host filesystem.
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
Yes, manual testing of the upgrade -> downgrade -> upgrade path validates that the Kubelet safely falls back to static boot-time caching without disrupting running workloads.
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?
No
Monitoring Requirements
Monitor the metrics
kubelet_node_resize_requests_totalkubelet_node_resize_errors_total
How can an operator determine if the feature is in use by workloads?
The enablement of the Kubelet feature gate can be determined via the kubernetes_feature_enabled metric. Operational use can be observed when the kubelet_node_resize_requests_total counter increments during a hardware change.
How can someone using this feature know that it is working for their instance?
An end-user can verify the feature by executing a hot-plug via their hypervisor, and then running kubectl get node <node-name> -o yaml. The .status.capacity and .status.allocatable fields will natively reflect the newly added hardware within ~10 seconds.
What are the reasonable SLOs (Service Level Objectives) for the enhancement?
For each dynamically resized node:
- Error rate: The
kubelet_node_resize_errors_totalcounter is expected to remain strictly at0during normal operations. - Latency (Spec-driven): For
Node.Spec.ConfiguredCapacity-triggered resize events, the time from Spec patch toCapacityConfigured: Acceptedcondition should complete within 30 seconds under normal conditions, bounded by informer propagation latency and cgroup write time. - Latency (Hardware-driven): For physical hardware events detected via cAdvisor polling, the reconciliation completes within one cAdvisor polling cycle (default: 5 minutes). Operators who require lower latency should configure a shorter cAdvisor polling interval.
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
- Metrics
- Metric name:
kubelet_node_resize_requests_totalkubelet_node_resize_errors_total
- Components exposing the metric: kubelet
- Metric name:
Are there any missing metrics that would be useful to have to improve observability of this feature?
No
Dependencies
Does this feature depend on any specific services running in the cluster?
cAdvisor (Internal): The Kubelet strictly relies on the integrated cAdvisor package to successfully read the underlying Linux kernel and hardware capacity.
Container Runtime (CRI): The Kubelet relies on the runtime (e.g., containerd, CRI-O) to successfully honor the UpdateContainerResources RPC call to propagate recalculated Swap limits.
API Server (Layer 2): The Node.Spec.ConfiguredCapacity field introduced in Layer 2 requires the API Server to recognise the updated NodeSpec schema. The API Server must be at a version that includes the new field before any Kubelet or external controller can use it.
Scheduler (Layer 1 validation + Layer 2): The Scheduler requires changes in two areas identified by this KEP: (a) a Node UPDATE queueing hint to re-queue pods in the WaitingForPreemption state when Node.Status.Allocatable decreases on their target node (needed for correctness with any mutable-capacity scenario, validated in Layer 1), and (b) using min(Node.Spec.ConfiguredCapacity, Node.Status.Capacity) as the effective capacity view during a downscale (Layer 2). If the Scheduler is not updated, the preemption grace period race and the scheduling-binding race window remain unmitigated.
Cluster Autoscaler (Layer 1 validation): The Layer 1 deliverable is to document and validate how CA behaves when Node.Status.Capacity changes — specifically the NodeGroup template corruption risk when a resized node is selected as the provisioning reference. The correct long-term fix (whether CA reads from the cloud provider’s launch template, or uses a configurable reference capacity in its own NodeGroup configuration) is an open design item that must be resolved before this KEP’s CA integration is standardised. CA changes are out of scope for this KEP’s feature gate.
Scalability
Will enabling / using this feature result in any new API calls?
Yes.
The Kubelet’s existing NodeInformer will now actively process and react to UPDATE events on the Node.Spec.ConfiguredCapacity field.
During a resize event, the Kubelet will issue PATCH calls to the Node status subresource to update .status.capacity, .status.allocatable, and .status.conditions. To prevent API thrashing during emergency hardware fluctuations, the Kubelet uses a jitter tolerance filter and caches the clamped state locally.
Will enabling / using this feature result in introducing new API types?
Yes. This introduces a new optional field ConfiguredCapacity within Node.Spec, and a new Node Condition type CapacityConfigured to track the reconciliation state (Accepted, InProgress, Infeasible, EmergencyReduced).
Will enabling / using this feature result in any new calls to the cloud provider?
No
Will enabling / using this feature result in increasing size or count of the existing API objects?
Yes.
- API type(s):
Node - Estimated increase in size: ~100-300 bytes per Node object. This is due to the addition of the
Node.Spec.ConfiguredCapacityfield and theCapacityConfiguredStatus Condition. No boot-time annotation or additional status field for the Autoscaler is introduced in the current design. - Estimated amount of new objects: 0 (No new objects are created).
Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?
Negligible, In the case of resource reconfiguration the resource manager may take some time to re-sync.
Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?
Negligible computational overhead is introduced. The Kubelet utilizes an event-driven Go channel to signal capacity updates, avoiding heavy CPU polling cycles.
Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?
Yes, organically. Expanding a node’s capacity allows the Scheduler to place more Pods onto the node, consuming more PIDs/sockets. However, this is strictly mitigated by the pre-existing --max-pods Kubelet configuration, which enforces a hard ceiling regardless of the underlying hardware size.
Troubleshooting
How does this feature react if the API server and/or etcd is unavailable?
If the API Server is unavailable during a hardware resize, the Kubelet will successfully update the local host cgroups, sub-managers, and Eviction thresholds to protect the node. However, the syncNodeStatus call will fail. The local node will be physically resized and stable, but the cluster Scheduler will remain “blind” to the new capacity until API connectivity is restored.
What are other known failure modes?
Cgroup Enforcement Failure
Detection: Spike in
kubelet_node_resize_errors_totalwith the labelsubsystem="cgroup_update".Mitigations: Investigate host filesystem or AppArmor/SELinux denials preventing the Kubelet from writing to
/sys/fs/cgroup.
CRI Swap Update Failure
Detection: Spike in
kubelet_node_resize_errors_totalwith the labelsubsystem="container_swap_resize".Mitigations: Verify the Container Runtime (containerd/CRI-O) is healthy and accepting RPC calls.
Downscale Stuck in InProgress (Eviction Stall)
Scenario: A downscale was initiated via
Node.Spec.ConfiguredCapacity. The Kubelet set theCapacityConfiguredcondition toFalse (Reason: InProgress)and began evicting starved pods. However, the eviction never completes — for example, because affected pods havePodDisruptionBudgetsthat block eviction, or the pods areGuaranteedQoS with no safe eviction path, or the Kubelet crashed mid-eviction.Detection: The
CapacityConfiguredcondition remainsFalse (Reason: InProgress)for longer than expected. Nokubelet_node_resize_requests_totalincrement is observed for the directiondecrease. The external controller’s watch on theAcceptedcondition never fires.Mitigations:
- Temporarily patch
Node.Spec.ConfiguredCapacityback to the previous (higher) value to cancel the downscale and unblock the node. The Kubelet will detect the spec revert, restore the cgroup boundaries, and transition the condition toAccepted. - Identify and resolve the blocking condition (e.g., adjust PodDisruptionBudgets, force-delete the stalled pod) and re-issue the downscale spec patch.
- As a last resort, disable the feature gate and restart the Kubelet to freeze capacity evaluation.
- Temporarily patch
What steps should be taken if SLOs are not being met to determine the problem?
Examine Kubelet logs for errors emitted by container_manager_linux.go. Disable the feature gate to freeze capacity evaluation until the host-level conflict is resolved.
Implementation History
- 2023-04-17: Initial KEP PR (#3955 ) opened as KEP-3953: Dynamic Node Resize — original scope covering both scale-up and scale-down via cAdvisor polling.
- 2024-01-31: Scope narrowed to scale-up only (Node Resource Hot Plug) following community feedback that a separate CRI-based hardware discovery mechanism was needed before scale-down could be safely addressed.
- 2025-01-13: KEP retitled to KEP-3953: Node Resource Hot Plug to reflect the updated focus on upscaling; Production Readiness Review Questionnaire updated.
- 2025-02-12: PRR approved for Alpha. Key design additions: swap limit recalculation for existing containers via
UpdateContainerResources, OOMScoreAdj drift accepted as a known limitation, hot-unplug emergency path outlined in Future Work. - 2026-02-10: KEP retitled to KEP-3953: In-place Node Resource Resize to reflect the full bidirectional resize scope introduced by the declarative
Node.Spec.ConfiguredCapacityAPI field — a major design pivot driven by reviewer feedback. - v1.38: Alpha milestone scoped to Layer 1 only (Ecosystem Tolerance). No new API fields or feature gate. The declarative Layer 2 design is documented in this KEP for community review but deferred to v1.39 pending resolution of the open design questions captured in Open Questions for Layer 2 .
Drawbacks
If dynamically removed resources (specifically CPUs or NUMA Memory zones) were exclusively pinned and allocated to specific containers via the CPU Manager or Topology Manager, the underlying hardware backing their strict isolation guarantees no longer exists on the motherboard.
Because the node can no longer fulfill the Pod’s strict Requests contract, those affected pods must be forcefully evicted by the Kubelet with a Failed status (Reason: NodeCapacityExceeded). While this protects the node, it does introduce a disruptive pod termination that would not occur if the node capacity remained static.
Alternatives
Node Replacement (Horizontal Scaling): Instead of resizing existing nodes, operators can horizontally scale by provisioning entirely new, larger nodes, draining the original nodes, and deleting them.
- Why it was rejected: This introduces significant workload disruption during the drain process, increases control plane overhead, and takes considerably longer to execute than a simple underlying hypervisor hot-plug.
Manual Kubelet Restarts: Administrators can hot-plug the hardware and then manually restart the
kubeletsystemd service to force it to read the new capacity.- Why it was rejected: This causes temporary node
NotReadystates, breaks activeexec/port-forwardsessions, and risks triggering historic edge-case bugs associated with Kubelet restarts.
- Why it was rejected: This causes temporary node
Balloon Drivers / Fake Placeholder Resources: Pre-provisioning massive virtual machines but using memory ballooning or fake placeholder resources to artificially restrict the Kubelet, “inflating” them when capacity is needed.
- Why it was rejected: This is highly inefficient, complex to manage at the hypervisor level, and confuses the Kubernetes Scheduler, which relies on accurate, native cgroup boundaries.
Node Annotation as Configuration Mechanism (
resize.node.kubernetes.io/configured-capacity): Using a Node annotation instead of a first-classNodeSpecfield to carry the desired capacity declaration.- Why it was rejected: Annotations are unstructured strings with no API validation, no admission webhook targeting support, and no defaulting semantics. They are effectively a workaround for the absence of a proper API field. Cluster administrators wanting to restrict who can set capacity would have to write brittle label-matching admission webhooks rather than using structured
ValidatingWebhookConfigurationfield selectors. A formalNodeSpecfield provides schema validation, cleankubectl diffoutput, and correct versioning/defaulting via the API machinery. An annotation is a hack around the right answer.
- Why it was rejected: Annotations are unstructured strings with no API validation, no admission webhook targeting support, and no defaulting semantics. They are effectively a workaround for the absence of a proper API field. Cluster administrators wanting to restrict who can set capacity would have to write brittle label-matching admission webhooks rather than using structured
Local Kubelet Configuration File (Static Config): Expressing the desired logical capacity via a local file on the node host (e.g., a
KubeletConfigurationfield), requiring a Kubelet restart to apply.- Why it was rejected: Local configuration fundamentally cannot be driven by external controllers. A cluster-level controller managing capacity across many nodes cannot atomically write a file to a remote node’s filesystem and then restart its Kubelet. This approach breaks the API-driven operational model of Kubernetes, makes orchestration of downscale workflows impossible without SSH/node access, and requires a Kubelet restart — defeating the core goal of this KEP.
CRI-Based Hardware Discovery (Alternative Trigger): Using a Container Runtime Interface (CRI) extension to deliver hardware capacity events to the Kubelet instead of relying on cAdvisor polling.
- Why it was not chosen for Alpha: A CRI-based discovery mechanism would require new CRI API additions and runtime support across containerd, CRI-O, and other runtimes — a multi-org coordination effort that is orthogonal to the Kubelet reconciliation logic this KEP introduces. The cAdvisor-based polling approach is available today on all supported runtimes and platforms. This alternative is tracked as a future evolution path via KEP-5224 (Node Resource Discovery) and is explicitly called out in the Future Work section.
Physical Hot-Unplug as a Considered-But-Deferred Trigger: Having the Kubelet proactively orchestrate or initiate physical hardware removal (i.e., calling a hypervisor API to perform hot-unplug) as part of a downscale flow.
- Why it was deferred: The Kubelet has no knowledge of the hypervisor or infrastructure layer. Introducing such a call would violate the single-responsibility principle and couple the Kubelet to provider-specific infrastructure APIs. The correct model is for an external controller (which does understand the infrastructure) to coordinate the physical hot-unplug after observing that the Kubelet has completed its graceful downscale (i.e.,
CapacityConfiguredcondition reachesAccepted). The KEP’s Path A2 flow explicitly documents this coordination contract.
- Why it was deferred: The Kubelet has no knowledge of the hypervisor or infrastructure layer. Introducing such a call would violate the single-responsibility principle and couple the Kubelet to provider-specific infrastructure APIs. The correct model is for an external controller (which does understand the infrastructure) to coordinate the physical hot-unplug after observing that the Kubelet has completed its graceful downscale (i.e.,
Open Questions for Layer 2
The following design questions remain open and must be resolved before Layer 2 implementation begins. They are recorded here so that the community can discuss them in parallel with the Layer 1 milestone.
Node.Spec.ConfiguredCapacityvs.NodeDesiredAllocatableThe current Layer 2 design proposes overwriting
Node.Status.CapacityviaNode.Spec.ConfiguredCapacity. An alternative approach is to keepNode.Status.Capacityanchored strictly to physical reality and instead introduce a separate declarative field — such asNodeDesiredAllocatable— that adjusts the scheduling bound without touching the raw capacity. This separation would eliminate several of the risks catalogued in the Risks and Mitigations section (e.g., API Status Clamping, cAdvisor polling latency) because the physical capacity field would remain immutable. The trade-off is a more complex mental model (two writable fields with different semantics) and potential ambiguity for components that currently treatCapacityandAllocatableas a single authoritative source. The Layer 1 milestone is expected to generate practical evidence that informs which approach is cleaner.Overcommit and Dense Burst Workloads
There is growing interest in supporting highly dense, bursty workloads (AI agents, serverless sandboxes) by restricting scheduling bounds (
NodeAllocatable) to maximise pod density while keeping parent cgroups intentionally wide, allowing concurrent short-lived workloads to burst into unallocated physical memory. This is in tension with the current Layer 2 design, which ties cgroup boundaries directly to the declared capacity. How the Kubelet should handle aConfiguredCapacityvalue that implies different cgroup ceiling vs. scheduler-visible capacity needs to be defined before Layer 2 can ship.Scheduler Race Window Mitigation
The current Layer 2 design introduces a race window: an external controller patches
Node.Spec.ConfiguredCapacity, and subsequently the Kubelet processes the event and changes bothCapacityandAllocatableinNode.Status, while the Scheduler only readsNode.Status.Allocatable. The proposed mitigation — having the Scheduler usemin(Node.Spec.ConfiguredCapacity, Node.Status.Capacity)as its effective capacity — requires a Scheduler change and needs community agreement on whether that change belongs in the Scheduler core or in a plugin.A related and distinct concern involves the preemption grace period: if node capacity decreases while a victim pod is in its termination grace period (following a preemption decision), the Scheduler’s original fit calculation for the preemptor pod is now stale. The correct fix is a Node UPDATE queueing hint that re-queues pods in the
WaitingForPreemptionstate whenNode.Status.Allocatabledecreases on their target node. Whether this hint already exists, or needs to be added, must be confirmed as part of the Layer 1 Scheduler contract work. Both themin()capacity view change and the queueing hint fix may ultimately be the same Scheduler change, or they may be independent — that needs to be determined before Layer 2 ships.Cluster Autoscaler NodeGroup Template Stability
The Cluster Autoscaler uses existing nodes as templates for provisioning new nodes within the same NodeGroup. When node capacity is mutable, a resized node may be selected as that template, causing newly provisioned nodes to be expected at the larger size while the cloud provider continues to provision at the original base size — resulting in scheduling failures. Three approaches have been considered:
- A
resize.node.kubernetes.io/initial-capacityannotation stamped by the Kubelet at first boot, read by CA as its reference baseline. This is low friction but annotations are brittle, lack schema validation, and are awkward when multiple actors read them. - A dedicated
Node.Status.InitialCapacityfield written once at first registration. More idiomatic, but introduces a new API field and conflates a historical boot value with live node status. - CA reading directly from the cloud provider’s launch template, or introducing a configurable reference capacity within CA’s own NodeGroup configuration. This decouples the problem from the Kubernetes Node object entirely, which is arguably the cleanest separation of concerns, but requires cloud-provider portability work outside this KEP.
None of these options is settled. An additional edge case complicates all three: if an entire NodeGroup has been uniformly resized, the “initial” capacity is no longer a meaningful reference and the NodeGroup has functionally changed size. The right approach must handle this case explicitly. This is tracked as a next-phase item to be explored after the Layer 1 milestone establishes the observed behaviour baseline.
- A
Infrastructure Needed (Optional)
For standard Kubernetes CI (e2e_node tests), no special infrastructure is needed because the tests will utilize a mocked cAdvisor client to simulate hardware capacity events.
However, for provider-specific end-to-end integration testing in the future, underlying infrastructure VMs that natively support CPU and Memory hot-plugging will be required to validate the complete hardware-to-API lifecycle.
Future Work
Dynamic System and Kube Reserved Adjustments
- Currently, the
--system-reservedand--kube-reservedvalues are static configurations defined during Kubelet bootstrap. If a node scales massively (e.g., from 16GB to 128GB of memory), the OS and Kubelet might organically require a dynamically scaled reservation rather than the original static threshold. Future iterations could explore allowing these reservations to be expressed as percentages or dynamic tunables.
- Currently, the
NRI (Node Resource Interface) Integration
- Extending the internal Kubelet resize event broadcaster so that external NRI plugins can natively subscribe to hardware capacity changes. This would allow third-party runtime wrappers and advanced topology managers to react to node upscales and downscales simultaneously with the Kubelet.
Event-Driven Hardware Detection (Node Resource Discovery )
- Currently, this KEP relies on
cAdvisorand lightweight host polling to detect physical hardware changes. In the future, as KEP-5224 matures, the responsibility of hardware discovery will shift toward the Container Runtime Interface (CRI) and external resource plugins. Once the CRI is capable of natively broadcasting dynamic hardware capacity events to the Kubelet, this KEP’s capacity reconciliation loop will be updated to subscribe directly to those CRI events.
- Currently, this KEP relies on
Node Capacity Overcommit (Logical > Physical)
- This KEP explicitly defers support for configuring
Node.Spec.ConfiguredCapacityto a value greater than the raw physical hardware capacity (e.g., reporting 48Gi of memory on a 32Gi machine backed by swap). This is a compelling use-case — particularly for swap-overcommit scenarios where the OS’s swap space provides a meaningful backing store for workloads that tolerate memory latency. However, enabling this in Alpha would break the Eviction Manager’s absolute threshold math and OOM-killer assumptions, which rely on the invariant that logical ≤ physical. Future work will define how the Eviction Manager, cgroup limits, and memory accounting interact when logical capacity exceeds physical, and will specify per-resource rules (e.g., Swap may be the first resource exempt from the Alpha overcommit restriction, while CPU and raw Memory remain bounded by physical reality).
- This KEP explicitly defers support for configuring
Cluster Autoscaler Integration
- Once the open design question in Open Questions for Layer 2 is resolved, the agreed approach will be implemented in a subsequent phase. The goal is to ensure CA-managed clusters can safely operate with nodes whose capacity changes over time, without risking NodeGroup template corruption or provisioning failures.