KEP-4438: Restarting sidecar containers during Pod termination

Implementation History
ALPHA Implementable
Created 2024-01-25
Updated 2026-09-07
Latest v1.38
Milestones
Alpha v1.38
Ownership
Owning SIG
SIG Node
Participating SIGs

KEP-4438: Restarting sidecar containers during Pod termination

Release Signoff Checklist

Items marked with (R) are required prior to targeting to a milestone / release.

  • (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
  • (R) KEP approvers have approved the KEP status as implementable
  • (R) Design details are appropriately documented
  • (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
    • e2e Tests for all Beta API Operations (endpoints)
    • (R) Ensure GA e2e tests meet requirements for Conformance Tests
    • (R) Minimum Two Week Window for GA e2e tests to prove flake free
  • (R) Graduation criteria is in place
  • (R) Production readiness review completed
  • (R) Production readiness review approved
  • “Implementation History” section is up-to-date for milestone
  • User-facing documentation has been created in kubernetes/website, for publication to kubernetes.io
  • Supporting documentation—e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes

Summary

Sidecar containers should be restarted if they exit prematurely when a Pod is terminated to ensure that they are running while the containers that should terminate prior to its termination are still running. This feature was originally planned for beta release in the initial KEP, but after further analysis, we decided to postpone and introduce a separate feature gate for it.

Motivation

The reason for this KEP is that restarting sidecar containers during Pod termination requires fundamental changes to the Pod lifecycle, and such changes cannot be introduced in beta (enabled by default) without risking disruption to users.

On the other hand, the code for sidecar containers is already well tested, and we are confident that many use cases will benefit from them even without the restart during Pod termination.

For this reason, we are introducing a separate feature gate for the sidecar containers KEP to decouple the two features and allow users to use sidecar containers without the refactoring required for the restart during Pod termination.

Goals

The following behaviors should be maintained during pod termination:

  • sidecar containers restarting
  • liveness, readiness and startup probing
  • container lifecycle hooks running
  • service account token rotation
  • secret and configmap volume updates

Non-Goals

Proposal

The proposal is to introduce a new feature gate for the sidecar containers KEP to decouple the sidecar feature from the restart during Pod termination feature and allow users to use sidecar containers without the refactoring required for the restart during Pod termination.

Please refer to the original KEP for the details of the sidecar containers feature: https://git.k8s.io/enhancements/keps/sig-node/753-sidecar-containers

User Stories (Optional)

Story 1

Story 2

Notes/Constraints/Caveats (Optional)

Alpha limitations

The proposed Alpha implementation targets v1.38 behind SidecarsRestartableDuringPodTermination, disabled by default. It restarts a previously started sidecar while application containers or later sidecars still need it, and stops the current instance at its ordered turn. The lifecycle design below is proposed for review alongside kubernetes/kubernetes#140133; that PR is not a merged implementation or evidence of design approval.

  • Probes: liveness and startup probes are stopped when termination begins. Probe workers are not reattached to replacement instances. Restarts are driven by observed exits, not probe failures. Probe support remains Beta work.
  • Recovery: API deletion deadlines survive kubelet restart through DeletionTimestamp. Local eviction and static-pod termination requests have no durable termination checkpoint; their original intent and deadline are not guaranteed to survive kubelet restart. See Recovery.
  • Hooks: postStart uses the normal start path. preStop is scheduled for each observed running instance, including replacements, while grace remains. Hook completion is not persisted, so a hook may execute again after kubelet restart. Hooks must tolerate replay.
  • Deadline expiry: no restart is attempted with at most one second remaining. At expiry, ordering and unfinished hooks no longer delay forced stops. The proposed zero-grace behavior differs from the legacy minimum stop grace and requires explicit review; see Deadlines and cancellation.
  • Availability: a deadline bounds the requested grace, not the time at which an unavailable runtime or hung node physically stops a process. Failed runtime observations keep termination pending and retain resources.
  • Configuration: pull secrets and image volumes are resolved through the normal start path. This does not add service-account token rotation or secret and configmap volume refresh during termination.
  • Observability: pod status is refreshed on each successful observation, including replacement IDs and restart counts. API publication remains asynchronous. The restart counter resets on kubelet restart and counts successful start-path completions, not starts whose CRI response was lost.

Pod termination and In-Place Pod Restart interaction

ShouldAllContainersRestart returns false for an API pod with a DeletionTimestamp. Once the worker enters TerminatingPod, it no longer calls SyncPod, including for local termination requests. Only the termination reconciler may restart eligible sidecars; it never starts application containers or triggers RestartAllContainers.

Risks and Mitigations

Changing SyncTerminatingPod from a one-shot operation to reconciliation changes when other kubelet subsystems may release resources. A successful RPC, an expired deadline, or a missing cache entry must not independently authorize cleanup. Tests exercise the worker completion signal, runtime observations, final status, and DRA unprepare boundary.

Lost CRI responses can leave a created or running replacement behind. A fresh runtime observation precedes each reconciliation, and created replacements are removed before retrying. Outstanding stop requests are deduplicated by container ID. Replays rely on CRI’s idempotent StopContainer and RemoveContainer contracts. Tests with the fake CRI cover failure before and after side effects; real-runtime node tests remain necessary to validate cancellation and ordering.

Restarting a sidecar continues using pod resources during termination, including images and credentials. The feature does not extend their validity or the pod’s grace budget. Alpha remains opt-in. Approval of the lifecycle decisions and passing node tests are release requirements, separate from unit-test success.

Design Details

Scope and worker transitions

The kubelet enables reconciliation only for a pod with restartable init containers when SidecarsRestartableDuringPodTermination is enabled. Ordinary pods, gate-disabled pods, sandbox replacement in SyncPod, and runtime-only orphan cleanup retain their existing kill paths. No API fields or CRI methods are added. The generic one-shot KillPod path does not contain a restart watcher.

The pod worker remains the sole owner of lifecycle transitions for a pod UID:

Current stateResultNext action
SyncPodTermination requested or normal execution finishedEnter TerminatingPod; stop normal setup
TerminatingPodcomplete=false, err=nilKeep resources and kill waiters; schedule another reconciliation
TerminatingPodErrorKeep resources and kill waiters; retry with backoff
TerminatingPodcomplete=true, err=nilNotify kill waiters and allow SyncTerminatedPod cleanup
TerminatedPodCleanup succeedsFinish the worker under the existing cleanup contract

The completion boolean is explicit: returning nil error does not mean the pod has stopped. An expired deadline also does not imply completion.

Observations and scheduling

Each invocation obtains the runtime pod and its container status directly, with one two-second context budget for the observation. The terminating worker skips podCache.GetNewerThan: that wait has no timeout and could otherwise prevent the worker’s retry timer from firing when PLEG stops advancing. A failed observation returns an error; it is not interpreted as an empty pod.

PLEG continues to supply ordinary observations and wakeups. It owns no desired termination state. The worker also owns a retry timer, normally one second for pending work. Errors use the existing worker backoff, capped by the time remaining until the pod deadline. After expiry, retries continue. The timer reuses the worker’s last pod specification, so eviction and removed static pods do not require another update from podManager. A shorter grace request cancels the current worker context and supplies an earlier deadline on the next invocation.

A slow start may occupy the worker until its context ends. Starts use the pod deadline and the worker cancellation context. The design depends on CRI and hook implementations honoring context cancellation; it does not promise progress through an indefinitely hung runtime call.

Desired actions and ordering

The reconciler derives desired actions from the pod specification, absolute deadline, and latest runtime observation:

  1. Application containers and non-restartable init containers are never started. Their observed live instances must stop.
  2. Walk restartable init containers in reverse specification order. A sidecar is still needed while any application container, non-restartable init container, or later sidecar is observed non-exited. Unknown state is conservatively live.
  3. A needed sidecar can restart only if it has previously started, is now exited (or has an unstarted replacement from a partial start), has a ready sandbox, and has more than one second left. Normal restart backoff applies. Missing status for a never-started sidecar does not authorize starting it.
  4. Once a sidecar’s turn arrives, stop its current observed instance. Do not restart a sidecar that exited at or after its turn. At the deadline, stop all remaining instances regardless of ordering and remove unstarted instances.

Starts use startContainer, including image pull secrets, image volumes and postStart. Secret and configmap managers are registered during termination so configuration can be resolved after kubelet restart. This registration does not restore volume-update or token-rotation behavior deferred from Alpha.

Runtime-manager records track outstanding hook and stop calls per container ID, and successful replacements per exited ID. They suppress duplicate work across repeated observations but do not define which containers should run. Calls complete through buffered channels; only the pod worker accesses these records.

Deadlines and cancellation

For a newly terminating worker, the local deadline is termination start plus the effective grace period. If the pod has a DeletionTimestamp, use the earlier of that timestamp and the local deadline. A shorter grace request can move the deadline earlier to request time plus the new grace. Repeated or longer requests never move it later within that worker’s lifetime.

preStop begins as soon as a running instance is observed during termination, including sidecars whose ordered stop is still pending. Each observed replacement gets its own hook. Hooks and ordering consume the same absolute grace budget. Hook completion is retained per instance for the lifetime of the runtime manager.

Hooks and stops run asynchronously so another reconciliation can observe exits and restart eligible sidecars. A stop passes the rounded-up remaining grace to CRI. Its RPC context allows two additional seconds for transport completion. When grace is shortened, a superseded stop context is cancelled and a new stop is issued for the same ID with the shorter grace. Cancellation is not evidence that the original server-side operation was rolled back or that the container stopped. Subsequent runtime observations determine progress.

At expiry, hooks are cancelled, new starts are forbidden, created instances are removed, and remaining containers receive a zero-grace stop. Each retry after expiry has a bounded stop RPC context. That transport allowance does not add container shutdown grace or authorize cleanup.

«[UNRESOLVED deadline compatibility]» The proposed implementation does not preserve the legacy kill path’s minimum two-second container grace after a long preStop or ordering wait. Reviewers must choose whether an absolute deadline should force immediately, as implemented, or whether a single bounded shutdown extension is required. If an extension is chosen, its recovery and shortening rules must be designed so retries cannot renew it. The deadline/hook tests currently assert zero-grace stops at expiry. «[/UNRESOLVED]»

Recovery

No new checkpoint is introduced in Alpha. On kubelet restart, spec and runtime status reconstruct ordering and restart eligibility. The API deletion timestamp reconstructs the original deadline even when it has already expired. Backoff and operation-deduplication records are in memory and may reset.

Interrupted operationObservation after restartRecovery
Create did not take effectPrevious sidecar exitedRetry through the normal start path
Create committed; start did notCreated replacement with a restart attemptRemove it and retry only while eligible; remove at expiry
Start committed; response lostRunning replacementKeep that instance; do not create another
Stop pending or response lostContainer still runningReissue idempotent stop using the remaining grace
Stop committedContainer exitedAdvance ordering without waiting for the old RPC result
preStop interrupted or completedSame instance still runningHook may replay while grace remains

These decisions assume the runtime reports committed operations and enforces container identity/name reservations for overlapping creation attempts. The in-memory records are not an exactly-once transaction across kubelet and CRI.

Local evictions and static-pod removals retain their deadline while the same worker exists. Without an API deletion timestamp, kubelet restart can lose the original local kill intent and deadline, as with existing local termination. If only a runtime pod remains and its spec is unavailable, the existing SyncTerminatingRuntimePod path stops it without sidecar restart.

«[UNRESOLVED local termination recovery]» Alpha proposes retaining the existing lack of durable local kill intent. Before claiming deadline preservation for all termination sources, design a checkpoint for both the intent and the absolute deadline, including eviction policy, static-pod replacement, and checkpoint cleanup. Persisting only a timestamp would not resolve recovery of the intent. Reviewers must explicitly accept this Alpha scope or require that work before Alpha. «[/UNRESOLVED]»

Completion, resources and status

The runtime reconciler returns pending while any observed container is non-exited, including unknown or created instances. Completed stop calls alone do not imply completion. Once observations show no remaining active containers, the kubelet stops the sandbox through the existing kill path, reads final runtime status, and checks that no containers remain running before unpreparing DRA resources and publishing final status. Only then may the worker transition to TerminatedPod and allow foreground cleanup and runtime removal. Existing status callbacks, such as eviction marking a pod Failed, may run before containers stop; API phase alone is not the authorization for resource removal.

Pod status is refreshed during reconciliation, so replacement IDs and restart counts are observable through the API subject to status publication latency. Liveness and startup probes stop when termination begins; all probe workers are removed after the pod stops. The kubelet_sidecar_restarts_during_termination_total counter increments after a successful restart path and resets when kubelet restarts.

Lifecycle invariants and review requirements

The implementation and tests must preserve these invariants:

  • No application-container restart, or sidecar restart after its turn or deadline.
  • Repeated observations and recovery of committed starts do not duplicate a live sidecar instance.
  • A worker’s deadline never extends; API deletion preserves it across kubelet restart. Local recovery is limited as described above.
  • PLEG silence does not block reconciliation. Runtime errors retain resources and schedule retries instead of reporting termination complete.
  • Stop response, timer expiry, and cancellation alone never release resources or kill waiters. Observed termination and final cleanup checks are required.
  • Gate-disabled pods, pods without sidecars, and runtime-only cleanup retain their existing lifecycle behavior.

The unresolved compatibility and recovery decisions require KEP approver review. Unit tests and this document do not imply that review has occurred. The node test suite must also run against a supported node/runtime before release.

Test Plan

[X] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.

Tests accompanying kubernetes/kubernetes#140133 exercise both the state machine and component boundaries. Fault-injection tests use a fake CRI; they do not replace execution of the node suite against a real runtime.

Prerequisite testing updates

The worker’s pending result, timer, and resource-removal predicates must be tested with the actual pod cache. A fake cache that always returns fresh status cannot expose a PLEG wait that blocks the deadline.

Unit tests
Invariant or failureTest in k8s.io/kubernetes
Stalled PLEG; API deletion, eviction and static-pod removal; shorter gracepkg/kubelet/pod_workers_termination_test.go: TestTerminatingPodProgressesWithStalledPLEG
Pending work retains waiters and runtime resourcesSame file: TestTerminatingPodRequeuesWithoutCompleting
Deadline monotonicity and expired API deadline recoverySame file: TestTerminationDeadlineDoesNotReset
New worker reconstructs an expired API deadlineSame file: TestTerminatingPodRecoversExpiredAPIDeadline
Runtime-only orphan cleanup never restarts sidecarsSame file: TestTerminatingRuntimePodDoesNotRestartSidecars
Error backoff cannot delay the next attempt beyond remaining graceSame file: TestTerminationRetryCannotPassDeadline
Fresh runtime observation supersedes stale cache; DRA and final-status guardpkg/kubelet/kubelet_termination_test.go: TestSyncTerminatingPodObservesRuntimeBeforeCleanup
Runtime observation failure remains pending and boundedSame file: TestSyncTerminatingPodObservationFailureRetainsResources
Gate controls the lifecycle pathSame file: TestSyncTerminatingPodGateControlsReconciliation
Lost create/start responses before or after side effects; new runtime managerpkg/kubelet/kuberuntime/kuberuntime_termination_restart_test.go: TestSyncTerminatingPodRecoversInterruptedStart
Stop completion or lost response is insufficient without an observationSame file: TestSyncTerminatingPodWaitsForObservedStop
Rejected stop is retried without renewing graceSame file: TestSyncTerminatingPodRetriesFailedStop
Shorter grace replaces an outstanding stop requestSame file: TestSyncTerminatingPodShortensOutstandingStop
Partial start cannot leave a created container past expirySame file: TestSyncTerminatingPodRemovesPartialStartAtDeadline
Earlier sidecar restarts while a later one drains; current instance stops in orderSame file: TestSyncTerminatingPodOrdersMultipleSidecars
Long hooks do not block reconciliation or renew graceSame file: TestSyncTerminatingPodPreStopDoesNotBlockReconciliation, TestSyncTerminatingPodDeadlineCancelsHooks
Restart eligibility, deduplication, backoff and retrySame file: TestSyncTerminatingPodDoesNotStartIneligibleSidecars, TestSyncTerminatingPodRestartsAndDeduplicates, TestSyncTerminatingPodRestartBackoff, TestSyncTerminatingPodRetriesPartialStart
Unknown container is stopped at expiry despite missing specSame file: TestSyncTerminatingPodDeadlineStopsUnknownContainer
No pod-wide restart during API deletionpkg/kubelet/container/helpers_test.go: TestShouldAllContainersRestart
Integration tests

The worker/cache and kubelet/runtime/resource-manager tests above run in package unit suites and exercise those component boundaries. No separate test/integration suite is claimed for this change. Real kubelet restart, API deletion and CRI process behavior are exercised by the node tests below.

e2e tests
Alpha implementation tests

test/e2e_node/sidecar_termination_restart_test.go, gated by SidecarsRestartableDuringPodTermination, contains these scenarios:

  • A sidecar exits during application shutdown, its restart count increases while the application is still running, and the pod subsequently terminates.
  • A later deletion with shorter grace overrides an ongoing termination; CRI observations confirm running containers disappear within the shortened budget.
  • Kubelet stops during termination, the original sidecar exits while it is down, and the restarted kubelet replaces it and completes within the original API deletion deadline (Serial, Disruptive).

These scenarios require execution on a supported node. Compilation and package fault-injection tests do not establish containerd or CRI-O behavior, nor prove that the disruptive kubelet-restart scenario passes.

Existing tests
Beta (planned)

The Alpha node scenarios above cover restart and grace-period behavior. Beta adds probe, hook replay, and configuration-lifetime assertions on real runtimes. Probe scenarios depend on implementing probe reattachment. Hook execution already uses the Alpha lifecycle paths; end-to-end validation must cover replacement instances and kubelet restart, not assume exactly-once delivery. Service-account token work (#116481, #122568) remains tracked separately.

Probes:

  • Readiness probes are still running while in preStop
  • Readiness status is beings updated for the container and the Pod while in preStop
  • Liveness probes are NOT running for regular containers while the Pod is terminating
  • SIDECAR: Liveness probes DO run for sidecar containers while the Pod is terminating
  • SIDECAR: sidecar container will be restarted when liveness probe failed during Pod termination

Not fully started containers:

  • preStop will not be executed for the container that hasn’t started yet
  • preStop will be called on the container even if postStart is still running
  • postStart hook CONTINUE EXECUTE even if container started termination
  • postStart hook will stop once pod passed it’s graceful termination period

Re-terminating the Pod:

  • When the Pod is terminating, another request with greater grace must not extend the deadline
  • BUGFIX: Service account token gets invalidated while terminating pod is re-deleted · Issue #122568

Pre-stop vs. SIGTERM traps:

  • Same as existing and above tests, need to validate that the container that traps the SIGTERM behaves the same way as with preStop:
  • Respect the grace period
  • Liveness probes are not running
  • Readiness probes are running

Test what is available for during preStop:

  • BUGFIX: While the Pod is terminating, service account tokens are rotated Kubelet stops rotating service account tokens when pod is terminating, breaking preStop hooks · Issue #116481
  • BUGFIX: Service account token is valid if the terminating Pod was deleted again Service account token gets invalidated while terminating pod is re-deleted · Issue #122568

Eviction and OOM kills:

  • preStop is called when Pod is evicted
  • preStop is NOT called when Container is OOMkilled

Graduation Criteria

Alpha

  • Feature implemented behind SidecarsRestartableDuringPodTermination, disabled by default, targeting v1.38.
  • KEP approvers accept the worker reconciliation contract and explicitly resolve the deadline-compatibility and local-recovery scope decisions above.
  • The lifecycle invariants have passing package tests, including fault injection at component boundaries and gate-disabled regression coverage.
  • Node restart, shorter-grace, and ordered shutdown scenarios execute successfully against a supported runtime. Test results are linked during implementation review; merely adding or compiling the tests is insufficient.

Beta

  • Resolve remaining Alpha limitations, including probe reattachment and its interaction with restart backoff during termination.
  • Decide whether to persist local termination intent/deadlines based on the Alpha recovery scope; test any durable recovery behavior before promising it.
  • Validate hook replay, replacement hooks, configuration lifetime, runtime cancellation and ambiguous CRI outcomes on real nodes.
  • Node tests pass in Testgrid without flakes for two consecutive releases.
  • Feedback from Alpha adopters is addressed.

GA

TBD

Upgrade / Downgrade Strategy

This feature only concerns the kubelet, so the upgrade and downgrade strategy is limited to the kubelet. Moreover, the Pod spec is not altered, so no changes are required for existing workloads to make use of the feature. Likewise, no changes are required for these workloads to revert to previous behavior.

Version Skew Strategy

There is no version skew strategy for this feature. The kubelet is the only component that needs to be updated to make use of this feature.

Production Readiness Review Questionnaire

Feature Enablement and Rollback

How can this feature be enabled / disabled in a live cluster?
  • Feature gate (also fill in values in kep.yaml)
    • Feature gate name: SidecarsRestartableDuringPodTermination
    • Components depending on the feature gate:
      • kubelet
  • Other
    • Describe the mechanism:
    • Will enabling / disabling the feature require downtime of the control plane?
    • Will enabling / disabling the feature require downtime or reprovisioning of a node?
Does enabling the feature change any default behavior?

Enabling the feature will change the behavior of the kubelet when terminating a Pod with sidecar containers. Sidecar containers that exit prematurely will be restarted during the termination of the Pod to ensure they are running until the main containers that should terminate prior to the sidecar containers are still running.

Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?

Yes, the feature can be disabled once it has been enabled. There is no alteration to the Pod spec, so existing workloads will be terminated according to the current behavior, after the kubelet is restarted with the feature gate disabled.

What happens if we reenable the feature if it was previously rolled back?

No side effect, the feature can be switched on or off.

Are there any tests for feature enablement/disablement?

Yes, unit tests will be added to ensure the feature can be enabled and disabled. The KEP will be updated with the details of the tests as they are added.

Rollout, Upgrade and Rollback Planning

How can a rollout or rollback fail? Can it impact already running workloads?
What specific metrics should inform a rollback?
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?

Monitoring Requirements

How can an operator determine if the feature is in use by workloads?

The kubelet_sidecar_restarts_during_termination_total counter (per node) is incremented every time a sidecar is restarted during pod termination. A non-zero and increasing value indicates the feature is enabled and actively restarting sidecars for terminating pods.

How can someone using this feature know that it is working for their instance?
  • Events
    • Event Reason:
  • API .status
    • Fields: initContainerStatuses[].containerID, restartCount, and state
    • Publication is asynchronous while the terminating pod still exists.
  • Other (treat as last resort)
    • Details:
What are the reasonable SLOs (Service Level Objectives) for the enhancement?
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
  • Metrics
    • Metric name: kubelet_sidecar_restarts_during_termination_total
    • Components exposing the metric: kubelet
  • Other (treat as last resort)
    • Details:
Are there any missing metrics that would be useful to have to improve observability of this feature?

Dependencies

Does this feature depend on any specific services running in the cluster?

Scalability

Will enabling / using this feature result in any new API calls?
Will enabling / using this feature result in introducing new API types?
Will enabling / using this feature result in any new calls to the cloud provider?
Will enabling / using this feature result in increasing size or count of the existing API objects?
Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?
Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?
Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?

Terminating sidecar pods add direct runtime status reads on each reconciliation. Pending timer retries normally occur once per second; pod updates can trigger additional reconciliations. This trades additional CRI reads for independence from stalled PLEG and stale observations after partial starts. Runtime latency and concurrent terminating-pod load need measurement before Beta.

Each pod retains one worker timer, per-instance operation records, and bounded hook/stop calls. Restart backoff limits repeated failing starts; deadline expiry forbids further starts. Replacements still consume normal pod resources, and the feature does not increase pod resource limits. A runtime that ignores cancellation can retain server-side work beyond a client timeout; real-runtime fault tests must cover this limitation.

Troubleshooting

How does this feature react if the API server and/or etcd is unavailable?
What are other known failure modes?
What steps should be taken if SLOs are not being met to determine the problem?

Implementation History

  • 2024-01-30: Summary and Motivation sections merged
  • 2024-02-08: Proposal section merged, KEP marked as implementable
  • 2026-09-07: Proposed worker reconciliation design and lifecycle regression tests in kubernetes/kubernetes#140133, targeting v1.38 Alpha. KEP lifecycle review and real-node validation remain release requirements.

Drawbacks

The main drawback of this KEP is that it introduces a new feature gate for the sidecars, which can be confusing for users. However, we believe that the current behavior of sidecar containers is already useful and that the restart during Pod termination feature is not critical for many use cases. This is why this feature is introduced as a separate feature gate, so that KEP-753 can reach GA faster.

Alternatives

The alternative would be to introduce KEP-753 with the restart during Pod termination feature. However, this would have required a significant refactoring of the kubelet and the Pod lifecycle, which would introduce a risk of disruption to users. This is why we decided to introduce a separate feature gate for the sidecar containers feature.

Infrastructure Needed (Optional)