KEP-4438: Restarting sidecar containers during Pod termination
KEP-4438: Restarting sidecar containers during Pod termination
- Release Signoff Checklist
- Summary
- Motivation
- Proposal
- Design Details
- Production Readiness Review Questionnaire
- Implementation History
- Drawbacks
- Alternatives
- Infrastructure Needed (Optional)
Release Signoff Checklist
Items marked with (R) are required prior to targeting to a milestone / release.
- (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
- (R) KEP approvers have approved the KEP status as
implementable - (R) Design details are appropriately documented
- (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
- e2e Tests for all Beta API Operations (endpoints)
- (R) Ensure GA e2e tests meet requirements for Conformance Tests
- (R) Minimum Two Week Window for GA e2e tests to prove flake free
- (R) Graduation criteria is in place
- (R) all GA Endpoints must be hit by Conformance Tests
- (R) Production readiness review completed
- (R) Production readiness review approved
- “Implementation History” section is up-to-date for milestone
- User-facing documentation has been created in kubernetes/website, for publication to kubernetes.io
- Supporting documentation—e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes
Summary
Sidecar containers should be restarted if they exit prematurely when a Pod is terminated to ensure that they are running while the containers that should terminate prior to its termination are still running. This feature was originally planned for beta release in the initial KEP, but after further analysis, we decided to postpone and introduce a separate feature gate for it.
Motivation
The reason for this KEP is that restarting sidecar containers during Pod termination requires fundamental changes to the Pod lifecycle, and such changes cannot be introduced in beta (enabled by default) without risking disruption to users.
On the other hand, the code for sidecar containers is already well tested, and we are confident that many use cases will benefit from them even without the restart during Pod termination.
For this reason, we are introducing a separate feature gate for the sidecar containers KEP to decouple the two features and allow users to use sidecar containers without the refactoring required for the restart during Pod termination.
Goals
The following behaviors should be maintained during pod termination:
- sidecar containers restarting
- liveness, readiness and startup probing
- container lifecycle hooks running
- service account token rotation
- secret and configmap volume updates
Non-Goals
Proposal
The proposal is to introduce a new feature gate for the sidecar containers KEP to decouple the sidecar feature from the restart during Pod termination feature and allow users to use sidecar containers without the refactoring required for the restart during Pod termination.
Please refer to the original KEP for the details of the sidecar containers feature: https://git.k8s.io/enhancements/keps/sig-node/753-sidecar-containers
User Stories (Optional)
Story 1
Story 2
Notes/Constraints/Caveats (Optional)
Alpha limitations
The proposed Alpha implementation targets v1.38 behind
SidecarsRestartableDuringPodTermination, disabled by default. It restarts a
previously started sidecar while application containers or later sidecars still
need it, and stops the current instance at its ordered turn. The lifecycle design
below is proposed for review alongside kubernetes/kubernetes#140133; that PR is
not a merged implementation or evidence of design approval.
- Probes: liveness and startup probes are stopped when termination begins. Probe workers are not reattached to replacement instances. Restarts are driven by observed exits, not probe failures. Probe support remains Beta work.
- Recovery: API deletion deadlines survive kubelet restart through
DeletionTimestamp. Local eviction and static-pod termination requests have no durable termination checkpoint; their original intent and deadline are not guaranteed to survive kubelet restart. See Recovery. - Hooks:
postStartuses the normal start path.preStopis scheduled for each observed running instance, including replacements, while grace remains. Hook completion is not persisted, so a hook may execute again after kubelet restart. Hooks must tolerate replay. - Deadline expiry: no restart is attempted with at most one second remaining. At expiry, ordering and unfinished hooks no longer delay forced stops. The proposed zero-grace behavior differs from the legacy minimum stop grace and requires explicit review; see Deadlines and cancellation.
- Availability: a deadline bounds the requested grace, not the time at which an unavailable runtime or hung node physically stops a process. Failed runtime observations keep termination pending and retain resources.
- Configuration: pull secrets and image volumes are resolved through the normal start path. This does not add service-account token rotation or secret and configmap volume refresh during termination.
- Observability: pod status is refreshed on each successful observation, including replacement IDs and restart counts. API publication remains asynchronous. The restart counter resets on kubelet restart and counts successful start-path completions, not starts whose CRI response was lost.
Pod termination and In-Place Pod Restart interaction
ShouldAllContainersRestart returns false for an API pod with a
DeletionTimestamp. Once the worker enters TerminatingPod, it no longer calls
SyncPod, including for local termination requests. Only the termination
reconciler may restart eligible sidecars; it never starts application containers
or triggers RestartAllContainers.
Risks and Mitigations
Changing SyncTerminatingPod from a one-shot operation to reconciliation changes
when other kubelet subsystems may release resources. A successful RPC, an expired
deadline, or a missing cache entry must not independently authorize cleanup.
Tests exercise the worker completion signal, runtime observations, final status,
and DRA unprepare boundary.
Lost CRI responses can leave a created or running replacement behind. A fresh
runtime observation precedes each reconciliation, and created replacements are
removed before retrying. Outstanding stop requests are deduplicated by container
ID. Replays rely on CRI’s idempotent StopContainer and RemoveContainer
contracts. Tests with the fake CRI cover failure before and after side effects;
real-runtime node tests remain necessary to validate cancellation and ordering.
Restarting a sidecar continues using pod resources during termination, including images and credentials. The feature does not extend their validity or the pod’s grace budget. Alpha remains opt-in. Approval of the lifecycle decisions and passing node tests are release requirements, separate from unit-test success.
Design Details
Scope and worker transitions
The kubelet enables reconciliation only for a pod with restartable init
containers when SidecarsRestartableDuringPodTermination is enabled. Ordinary
pods, gate-disabled pods, sandbox replacement in SyncPod, and runtime-only
orphan cleanup retain their existing kill paths. No API fields or CRI methods
are added. The generic one-shot KillPod path does not contain a restart watcher.
The pod worker remains the sole owner of lifecycle transitions for a pod UID:
| Current state | Result | Next action |
|---|---|---|
SyncPod | Termination requested or normal execution finished | Enter TerminatingPod; stop normal setup |
TerminatingPod | complete=false, err=nil | Keep resources and kill waiters; schedule another reconciliation |
TerminatingPod | Error | Keep resources and kill waiters; retry with backoff |
TerminatingPod | complete=true, err=nil | Notify kill waiters and allow SyncTerminatedPod cleanup |
TerminatedPod | Cleanup succeeds | Finish the worker under the existing cleanup contract |
The completion boolean is explicit: returning nil error does not mean the pod has stopped. An expired deadline also does not imply completion.
Observations and scheduling
Each invocation obtains the runtime pod and its container status directly, with
one two-second context budget for the observation. The terminating worker skips
podCache.GetNewerThan: that wait has no timeout and could otherwise prevent the
worker’s retry timer from firing when PLEG stops advancing. A failed observation
returns an error; it is not interpreted as an empty pod.
PLEG continues to supply ordinary observations and wakeups. It owns no desired termination state. The worker also owns a retry timer, normally one second for pending work. Errors use the existing worker backoff, capped by the time remaining until the pod deadline. After expiry, retries continue. The timer reuses the worker’s last pod specification, so eviction and removed static pods do not require another update from podManager. A shorter grace request cancels the current worker context and supplies an earlier deadline on the next invocation.
A slow start may occupy the worker until its context ends. Starts use the pod deadline and the worker cancellation context. The design depends on CRI and hook implementations honoring context cancellation; it does not promise progress through an indefinitely hung runtime call.
Desired actions and ordering
The reconciler derives desired actions from the pod specification, absolute deadline, and latest runtime observation:
- Application containers and non-restartable init containers are never started. Their observed live instances must stop.
- Walk restartable init containers in reverse specification order. A sidecar is still needed while any application container, non-restartable init container, or later sidecar is observed non-exited. Unknown state is conservatively live.
- A needed sidecar can restart only if it has previously started, is now exited (or has an unstarted replacement from a partial start), has a ready sandbox, and has more than one second left. Normal restart backoff applies. Missing status for a never-started sidecar does not authorize starting it.
- Once a sidecar’s turn arrives, stop its current observed instance. Do not restart a sidecar that exited at or after its turn. At the deadline, stop all remaining instances regardless of ordering and remove unstarted instances.
Starts use startContainer, including image pull secrets, image volumes and
postStart. Secret and configmap managers are registered during termination so
configuration can be resolved after kubelet restart. This registration does not
restore volume-update or token-rotation behavior deferred from Alpha.
Runtime-manager records track outstanding hook and stop calls per container ID, and successful replacements per exited ID. They suppress duplicate work across repeated observations but do not define which containers should run. Calls complete through buffered channels; only the pod worker accesses these records.
Deadlines and cancellation
For a newly terminating worker, the local deadline is termination start plus the
effective grace period. If the pod has a DeletionTimestamp, use the earlier of
that timestamp and the local deadline. A shorter grace request can move the
deadline earlier to request time plus the new grace. Repeated or longer requests
never move it later within that worker’s lifetime.
preStop begins as soon as a running instance is observed during termination,
including sidecars whose ordered stop is still pending. Each observed replacement
gets its own hook. Hooks and ordering consume the same absolute grace budget.
Hook completion is retained per instance for the lifetime of the runtime manager.
Hooks and stops run asynchronously so another reconciliation can observe exits and restart eligible sidecars. A stop passes the rounded-up remaining grace to CRI. Its RPC context allows two additional seconds for transport completion. When grace is shortened, a superseded stop context is cancelled and a new stop is issued for the same ID with the shorter grace. Cancellation is not evidence that the original server-side operation was rolled back or that the container stopped. Subsequent runtime observations determine progress.
At expiry, hooks are cancelled, new starts are forbidden, created instances are removed, and remaining containers receive a zero-grace stop. Each retry after expiry has a bounded stop RPC context. That transport allowance does not add container shutdown grace or authorize cleanup.
«[UNRESOLVED deadline compatibility]»
The proposed implementation does not preserve the legacy kill path’s minimum
two-second container grace after a long preStop or ordering wait. Reviewers must
choose whether an absolute deadline should force immediately, as implemented, or
whether a single bounded shutdown extension is required. If an extension is
chosen, its recovery and shortening rules must be designed so retries cannot
renew it. The deadline/hook tests currently assert zero-grace stops at expiry.
«[/UNRESOLVED]»
Recovery
No new checkpoint is introduced in Alpha. On kubelet restart, spec and runtime status reconstruct ordering and restart eligibility. The API deletion timestamp reconstructs the original deadline even when it has already expired. Backoff and operation-deduplication records are in memory and may reset.
| Interrupted operation | Observation after restart | Recovery |
|---|---|---|
| Create did not take effect | Previous sidecar exited | Retry through the normal start path |
| Create committed; start did not | Created replacement with a restart attempt | Remove it and retry only while eligible; remove at expiry |
| Start committed; response lost | Running replacement | Keep that instance; do not create another |
| Stop pending or response lost | Container still running | Reissue idempotent stop using the remaining grace |
| Stop committed | Container exited | Advance ordering without waiting for the old RPC result |
preStop interrupted or completed | Same instance still running | Hook may replay while grace remains |
These decisions assume the runtime reports committed operations and enforces container identity/name reservations for overlapping creation attempts. The in-memory records are not an exactly-once transaction across kubelet and CRI.
Local evictions and static-pod removals retain their deadline while the same
worker exists. Without an API deletion timestamp, kubelet restart can lose the
original local kill intent and deadline, as with existing local termination.
If only a runtime pod remains and its spec is unavailable, the existing
SyncTerminatingRuntimePod path stops it without sidecar restart.
«[UNRESOLVED local termination recovery]» Alpha proposes retaining the existing lack of durable local kill intent. Before claiming deadline preservation for all termination sources, design a checkpoint for both the intent and the absolute deadline, including eviction policy, static-pod replacement, and checkpoint cleanup. Persisting only a timestamp would not resolve recovery of the intent. Reviewers must explicitly accept this Alpha scope or require that work before Alpha. «[/UNRESOLVED]»
Completion, resources and status
The runtime reconciler returns pending while any observed container is non-exited,
including unknown or created instances. Completed stop calls alone do not imply
completion. Once observations show no remaining active containers, the kubelet
stops the sandbox through the existing kill path, reads final runtime status,
and checks that no containers remain running before unpreparing DRA resources
and publishing final status. Only then may the worker transition to
TerminatedPod and allow foreground cleanup and runtime removal. Existing status
callbacks, such as eviction marking a pod Failed, may run before containers stop;
API phase alone is not the authorization for resource removal.
Pod status is refreshed during reconciliation, so replacement IDs and restart
counts are observable through the API subject to status publication latency.
Liveness and startup probes stop when termination begins; all probe workers are
removed after the pod stops. The
kubelet_sidecar_restarts_during_termination_total counter increments after a
successful restart path and resets when kubelet restarts.
Lifecycle invariants and review requirements
The implementation and tests must preserve these invariants:
- No application-container restart, or sidecar restart after its turn or deadline.
- Repeated observations and recovery of committed starts do not duplicate a live sidecar instance.
- A worker’s deadline never extends; API deletion preserves it across kubelet restart. Local recovery is limited as described above.
- PLEG silence does not block reconciliation. Runtime errors retain resources and schedule retries instead of reporting termination complete.
- Stop response, timer expiry, and cancellation alone never release resources or kill waiters. Observed termination and final cleanup checks are required.
- Gate-disabled pods, pods without sidecars, and runtime-only cleanup retain their existing lifecycle behavior.
The unresolved compatibility and recovery decisions require KEP approver review. Unit tests and this document do not imply that review has occurred. The node test suite must also run against a supported node/runtime before release.
Test Plan
[X] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.
Tests accompanying kubernetes/kubernetes#140133 exercise both the state machine and component boundaries. Fault-injection tests use a fake CRI; they do not replace execution of the node suite against a real runtime.
Prerequisite testing updates
The worker’s pending result, timer, and resource-removal predicates must be tested with the actual pod cache. A fake cache that always returns fresh status cannot expose a PLEG wait that blocks the deadline.
Unit tests
| Invariant or failure | Test in k8s.io/kubernetes |
|---|---|
| Stalled PLEG; API deletion, eviction and static-pod removal; shorter grace | pkg/kubelet/pod_workers_termination_test.go: TestTerminatingPodProgressesWithStalledPLEG |
| Pending work retains waiters and runtime resources | Same file: TestTerminatingPodRequeuesWithoutCompleting |
| Deadline monotonicity and expired API deadline recovery | Same file: TestTerminationDeadlineDoesNotReset |
| New worker reconstructs an expired API deadline | Same file: TestTerminatingPodRecoversExpiredAPIDeadline |
| Runtime-only orphan cleanup never restarts sidecars | Same file: TestTerminatingRuntimePodDoesNotRestartSidecars |
| Error backoff cannot delay the next attempt beyond remaining grace | Same file: TestTerminationRetryCannotPassDeadline |
| Fresh runtime observation supersedes stale cache; DRA and final-status guard | pkg/kubelet/kubelet_termination_test.go: TestSyncTerminatingPodObservesRuntimeBeforeCleanup |
| Runtime observation failure remains pending and bounded | Same file: TestSyncTerminatingPodObservationFailureRetainsResources |
| Gate controls the lifecycle path | Same file: TestSyncTerminatingPodGateControlsReconciliation |
| Lost create/start responses before or after side effects; new runtime manager | pkg/kubelet/kuberuntime/kuberuntime_termination_restart_test.go: TestSyncTerminatingPodRecoversInterruptedStart |
| Stop completion or lost response is insufficient without an observation | Same file: TestSyncTerminatingPodWaitsForObservedStop |
| Rejected stop is retried without renewing grace | Same file: TestSyncTerminatingPodRetriesFailedStop |
| Shorter grace replaces an outstanding stop request | Same file: TestSyncTerminatingPodShortensOutstandingStop |
| Partial start cannot leave a created container past expiry | Same file: TestSyncTerminatingPodRemovesPartialStartAtDeadline |
| Earlier sidecar restarts while a later one drains; current instance stops in order | Same file: TestSyncTerminatingPodOrdersMultipleSidecars |
| Long hooks do not block reconciliation or renew grace | Same file: TestSyncTerminatingPodPreStopDoesNotBlockReconciliation, TestSyncTerminatingPodDeadlineCancelsHooks |
| Restart eligibility, deduplication, backoff and retry | Same file: TestSyncTerminatingPodDoesNotStartIneligibleSidecars, TestSyncTerminatingPodRestartsAndDeduplicates, TestSyncTerminatingPodRestartBackoff, TestSyncTerminatingPodRetriesPartialStart |
| Unknown container is stopped at expiry despite missing spec | Same file: TestSyncTerminatingPodDeadlineStopsUnknownContainer |
| No pod-wide restart during API deletion | pkg/kubelet/container/helpers_test.go: TestShouldAllContainersRestart |
Integration tests
The worker/cache and kubelet/runtime/resource-manager tests above run in package
unit suites and exercise those component boundaries. No separate
test/integration suite is claimed for this change. Real kubelet restart,
API deletion and CRI process behavior are exercised by the node tests below.
e2e tests
Alpha implementation tests
test/e2e_node/sidecar_termination_restart_test.go, gated by
SidecarsRestartableDuringPodTermination, contains these scenarios:
- A sidecar exits during application shutdown, its restart count increases while the application is still running, and the pod subsequently terminates.
- A later deletion with shorter grace overrides an ongoing termination; CRI observations confirm running containers disappear within the shortened budget.
- Kubelet stops during termination, the original sidecar exits while it is down,
and the restarted kubelet replaces it and completes within the original API
deletion deadline (
Serial,Disruptive).
These scenarios require execution on a supported node. Compilation and package fault-injection tests do not establish containerd or CRI-O behavior, nor prove that the disruptive kubelet-restart scenario passes.
Existing tests
- should respect termination grace period seconds
- should respect termination grace period seconds with long-running preStop hook https://github.com/kubernetes/kubernetes/blob/fbb2e6293fb0c8c107ae48b8b8ae488325c59598/test/e2e_node/container_lifecycle_test.go#L536
- should call the container’s preStop hook and terminate it if its startup probe fails https://github.com/kubernetes/kubernetes/blob/master/test/e2e_node/container_lifecycle_test.go#L616
- should call the container’s preStop hook and terminate it if its liveness probe fails https://github.com/kubernetes/kubernetes/blob/fbb2e6293fb0c8c107ae48b8b8ae488325c59598/test/e2e_node/container_lifecycle_test.go#L683
Beta (planned)
The Alpha node scenarios above cover restart and grace-period behavior. Beta adds probe, hook replay, and configuration-lifetime assertions on real runtimes. Probe scenarios depend on implementing probe reattachment. Hook execution already uses the Alpha lifecycle paths; end-to-end validation must cover replacement instances and kubelet restart, not assume exactly-once delivery. Service-account token work (#116481, #122568) remains tracked separately.
Probes:
- Readiness probes are still running while in preStop
- Readiness status is beings updated for the container and the Pod while in preStop
- Liveness probes are NOT running for regular containers while the Pod is terminating
- SIDECAR: Liveness probes DO run for sidecar containers while the Pod is terminating
- SIDECAR: sidecar container will be restarted when liveness probe failed during Pod termination
Not fully started containers:
- preStop will not be executed for the container that hasn’t started yet
- preStop will be called on the container even if postStart is still running
- postStart hook CONTINUE EXECUTE even if container started termination
- postStart hook will stop once pod passed it’s graceful termination period
Re-terminating the Pod:
- When the Pod is terminating, another request with greater grace must not extend the deadline
- BUGFIX: Service account token gets invalidated while terminating pod is re-deleted · Issue #122568
Pre-stop vs. SIGTERM traps:
- Same as existing and above tests, need to validate that the container that traps the SIGTERM behaves the same way as with preStop:
- Respect the grace period
- Liveness probes are not running
- Readiness probes are running
Test what is available for during preStop:
- BUGFIX: While the Pod is terminating, service account tokens are rotated Kubelet stops rotating service account tokens when pod is terminating, breaking preStop hooks · Issue #116481
- BUGFIX: Service account token is valid if the terminating Pod was deleted again Service account token gets invalidated while terminating pod is re-deleted · Issue #122568
Eviction and OOM kills:
- preStop is called when Pod is evicted
- preStop is NOT called when Container is OOMkilled
Graduation Criteria
Alpha
- Feature implemented behind
SidecarsRestartableDuringPodTermination, disabled by default, targeting v1.38. - KEP approvers accept the worker reconciliation contract and explicitly resolve the deadline-compatibility and local-recovery scope decisions above.
- The lifecycle invariants have passing package tests, including fault injection at component boundaries and gate-disabled regression coverage.
- Node restart, shorter-grace, and ordered shutdown scenarios execute successfully against a supported runtime. Test results are linked during implementation review; merely adding or compiling the tests is insufficient.
Beta
- Resolve remaining Alpha limitations, including probe reattachment and its interaction with restart backoff during termination.
- Decide whether to persist local termination intent/deadlines based on the Alpha recovery scope; test any durable recovery behavior before promising it.
- Validate hook replay, replacement hooks, configuration lifetime, runtime cancellation and ambiguous CRI outcomes on real nodes.
- Node tests pass in Testgrid without flakes for two consecutive releases.
- Feedback from Alpha adopters is addressed.
GA
TBD
Upgrade / Downgrade Strategy
This feature only concerns the kubelet, so the upgrade and downgrade strategy is limited to the kubelet. Moreover, the Pod spec is not altered, so no changes are required for existing workloads to make use of the feature. Likewise, no changes are required for these workloads to revert to previous behavior.
Version Skew Strategy
There is no version skew strategy for this feature. The kubelet is the only component that needs to be updated to make use of this feature.
Production Readiness Review Questionnaire
Feature Enablement and Rollback
How can this feature be enabled / disabled in a live cluster?
- Feature gate (also fill in values in
kep.yaml)- Feature gate name: SidecarsRestartableDuringPodTermination
- Components depending on the feature gate:
- kubelet
- Other
- Describe the mechanism:
- Will enabling / disabling the feature require downtime of the control plane?
- Will enabling / disabling the feature require downtime or reprovisioning of a node?
Does enabling the feature change any default behavior?
Enabling the feature will change the behavior of the kubelet when terminating a Pod with sidecar containers. Sidecar containers that exit prematurely will be restarted during the termination of the Pod to ensure they are running until the main containers that should terminate prior to the sidecar containers are still running.
Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?
Yes, the feature can be disabled once it has been enabled. There is no alteration to the Pod spec, so existing workloads will be terminated according to the current behavior, after the kubelet is restarted with the feature gate disabled.
What happens if we reenable the feature if it was previously rolled back?
No side effect, the feature can be switched on or off.
Are there any tests for feature enablement/disablement?
Yes, unit tests will be added to ensure the feature can be enabled and disabled. The KEP will be updated with the details of the tests as they are added.
Rollout, Upgrade and Rollback Planning
How can a rollout or rollback fail? Can it impact already running workloads?
What specific metrics should inform a rollback?
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?
Monitoring Requirements
How can an operator determine if the feature is in use by workloads?
The kubelet_sidecar_restarts_during_termination_total counter (per node) is incremented every
time a sidecar is restarted during pod termination. A non-zero and increasing value indicates
the feature is enabled and actively restarting sidecars for terminating pods.
How can someone using this feature know that it is working for their instance?
- Events
- Event Reason:
- API .status
- Fields:
initContainerStatuses[].containerID,restartCount, andstate - Publication is asynchronous while the terminating pod still exists.
- Fields:
- Other (treat as last resort)
- Details:
What are the reasonable SLOs (Service Level Objectives) for the enhancement?
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
- Metrics
- Metric name:
kubelet_sidecar_restarts_during_termination_total - Components exposing the metric: kubelet
- Metric name:
- Other (treat as last resort)
- Details:
Are there any missing metrics that would be useful to have to improve observability of this feature?
Dependencies
Does this feature depend on any specific services running in the cluster?
Scalability
Will enabling / using this feature result in any new API calls?
Will enabling / using this feature result in introducing new API types?
Will enabling / using this feature result in any new calls to the cloud provider?
Will enabling / using this feature result in increasing size or count of the existing API objects?
Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?
Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?
Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?
Terminating sidecar pods add direct runtime status reads on each reconciliation. Pending timer retries normally occur once per second; pod updates can trigger additional reconciliations. This trades additional CRI reads for independence from stalled PLEG and stale observations after partial starts. Runtime latency and concurrent terminating-pod load need measurement before Beta.
Each pod retains one worker timer, per-instance operation records, and bounded hook/stop calls. Restart backoff limits repeated failing starts; deadline expiry forbids further starts. Replacements still consume normal pod resources, and the feature does not increase pod resource limits. A runtime that ignores cancellation can retain server-side work beyond a client timeout; real-runtime fault tests must cover this limitation.
Troubleshooting
How does this feature react if the API server and/or etcd is unavailable?
What are other known failure modes?
What steps should be taken if SLOs are not being met to determine the problem?
Implementation History
- 2024-01-30:
SummaryandMotivationsections merged - 2024-02-08:
Proposalsection merged, KEP marked asimplementable - 2026-09-07: Proposed worker reconciliation design and lifecycle regression tests in kubernetes/kubernetes#140133, targeting v1.38 Alpha. KEP lifecycle review and real-node validation remain release requirements.
Drawbacks
The main drawback of this KEP is that it introduces a new feature gate for the sidecars, which can be confusing for users. However, we believe that the current behavior of sidecar containers is already useful and that the restart during Pod termination feature is not critical for many use cases. This is why this feature is introduced as a separate feature gate, so that KEP-753 can reach GA faster.
Alternatives
The alternative would be to introduce KEP-753 with the restart during Pod termination feature. However, this would have required a significant refactoring of the kubelet and the Pod lifecycle, which would introduce a risk of disruption to users. This is why we decided to introduce a separate feature gate for the sidecar containers feature.