KEP-6404: API Server Write Throughput: Reducing Allocations
KEP-6404: API Server Write Throughput: Reducing Allocations
- Release Signoff Checklist
- Summary
- Motivation
- Proposal
- Design Details
- Production Readiness Review Questionnaire
- Implementation History
- Drawbacks
- Alternatives
- Potential Future Extensions
- Infrastructure Needed (Optional)
Release Signoff Checklist
Items marked with (R) are required prior to targeting to a milestone / release.
- (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
- (R) KEP approvers have approved the KEP status as
implementable - (R) Design details are appropriately documented
- (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
- e2e Tests for all Beta API Operations (endpoints)
- (R) Ensure GA e2e tests meet requirements for Conformance Tests
- (R) Minimum Two Week Window for GA e2e tests to prove flake free
- (R) Graduation criteria is in place
- (R) all GA Endpoints must be hit by Conformance Tests within one minor version of promotion to GA
- (R) Production readiness review completed
- (R) Production readiness review approved
- “Implementation History” section is up-to-date for milestone
- User-facing documentation has been created in kubernetes/website , for publication to kubernetes.io
- Supporting documentation—e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes
Summary
Following recent SIG Scalability investigations into kube-apiserver write throughput using realistic resource sizes (>10KB pods), memory allocations were identified as the primary bottleneck limiting write scale and causing sticky performance degradation during disruptions.
This KEP proposes to address both sources of allocations on happy path and remove some kube-apiserver fallbacks that lead to allocations amplifaction causing unrecoverable state degradations.
Motivation
Following the SIG Scalability proposals to introduce Resource Size as a new dimension (kubernetes/kubernetes#134375
) and establish a realistic “Pod Shape” (kubernetes/kubernetes#138415
), write throughput benchmarks revealed that memory allocations are the primary limiting factor for kube-apiserver scalability.
1. The Allocation Ceiling
The API server allocates over 1MB of memory to apply a trivial patch on a 10KB pod object—a >100x allocation factor on the happy path (which includes storing the object in etcd, observing it on watch, decoding and transforming it, storing it in the watch cache, and re-encoding/propagating it to external and internal watchers).
When running k8s write throughput benchmark
on a c4-highmem-144 machine, we discovered that Go runtime hits a hard ceiling around 6GB/s of memory allocations. At this ceiling, Compare-And-Swap (CAS) operations during the GC sweep phase consume >50% of CPU, saturating the node. While tuning GOGC provides temporary relief (up to ~12GB/s), while we reported the issue to Go team, a fundamental reduction in per-operation allocations within kube-apiserver is required.
Sources of these allocations include:
- Inefficient PATCH validation and merging: Specifically
managedFieldspatching and Server-Side Apply (SSA) tracking, which account for over half of write-path allocations (tracked separately in kubernetes/kubernetes#142228 ). - Repeated Whole-Object Decoding and Encoding: Even when a request only modifies tens of bytes on a 10KB object (for example, a Pod
PATCHthat setsspec.nodeNameor updatesstatus),kube-apiserveroperates on the object as a whole rather than only on the relevant bytes that changed. While the watch cache helps serialize objects once for synced watchers, each operation on an object serializes the entire object from scratch.
While managedFields overhead is being optimized directly in kubernetes/kubernetes#142228
, avoiding redundant whole-object encoding and decoding has nontrivial interactions with Kubernetes’s data model (defaulting, unknown field stripping, and storage decoration) and is documented in Potential Future Extensions
.
2. Sticky Degradation on Disruptions
The >100x allocation factor only accounts for the steady-state happy path. On a traffic burst or short disruption, the API server can enter a degraded state where the fallback path is more costly than the normal path, preventing recovery. Common causes of such a state include:
- Watch breaking: When a watch connection breaks, re-establishing watch on non-zero RV doesn’t benefit from serialization caching.
- Admission Informer: The API server includes built-in informers that on short disruption can relist, further straining the process.
- Conflicting TXNs: When optimistic concurrency conflicts occur, the API server falls back to fetching the latest object directly from etcd and decoding it from scratch.
In these scenarios, allocations easily double, turning a brief disruption into a persistent, sticky degradation. Hitting the allocation ceiling requires dropping client throughput by 2x–3x before kube-apiserver can recover. Consequently, kube-apiserver cannot tolerate sudden bursts in write QPS without entering a degraded loop.
As part of this KEP we would like to tackle the problem of conflicting TXN fallbacks by serving all GETs from the watch cache.
Goals
- Reduce per-operation memory allocations in
kube-apiserverto raise the write-throughput ceiling and prevent sticky performance degradation during disruptions. - Serve all
GETrequests from the watch cache, including consistentGETs (resourceVersion="") and internalGETs on write/patch transaction (TXN) conflicts.
Non-Goals
Proposal
Serve All GETs from the Watch Cache
Currently, while LIST requests with resourceVersion="" use the consistent read-from-cache mechanism (KEP-2340
) and historical LISTs use B-tree cache snapshots (KEP-4988
), many GET paths still bypass the watch cache and hit etcd directly:
- Consistent
GETrequests (resourceVersion=""). - Internal
GETrequests executed byGuaranteedUpdatewhen an optimistic concurrency transaction (TXN) fails due to a conflict.
We propose routing all GET requests through the watch cache:
- For consistent
GETs andTXNconflict retries, the watch cache will ensure freshness up to the required etcd revision (using watch progress notifications or the revision returned by the failed TXN response) and return the already-decoded cached object from the immutable B-tree snapshot. - Because the watch cache in v1.37+ relies on immutable B-tree snapshots and avoids copying under the lock, serving all
GETs from cache scales concurrently without lock contention and eliminates both etcd load spikes and expensive per-retry decode allocations during write conflicts.
Risks and Mitigations
Watch Cache Lock Contention
- Risk: Routing all
GETs (includingGuaranteedUpdateconflictGETs) through the watch cache could increase lock contention on the watch cache mutex. - Mitigation: Major watch cache refactors in v1.37 replaced copying under locks with lock-free reads over immutable B-tree snapshots.
GETlookups only acquire a brief pointer to the latest snapshot.
Watch Cache Correctness
- Risk: Ensuring watch cache can provide same consistency guarantees as etcd has proven to be a risk in past efforts, there is a risk for subtile bugs that are very hard to detect with traditional testing.
- Mitigation: We will introduce mode-based correctness and linearizability testing for k8s storage as proposed in https://github.com/kubernetes/kubernetes/issues/141652 .
Design Details
- Serving All
GETs from Cache (ConsistentGetFromCache):- For client consistent
GETs (resourceVersion=""),cacher.Getdetermines the required revision from etcd (or uses the latest known revision) and invokesWaitUntilFreshAndGet(ctx, requiredRV, key)against the watch cache B-tree snapshot. - For optimistic concurrency (
TXN) conflicts inGuaranteedUpdate, instead of fetching and decoding the conflicting object directly from etcd,GuaranteedUpdatewaits for the watch cache to reach the revision returned in the failedTxnResponseand retrieves the already-decoded object from the watch cache.
- For client consistent
Test Plan
[x] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.
Prerequisite testing updates
None.
Unit tests
- Unit tests in
k8s.io/apiserver/pkg/storage/cacherandk8s.io/apiserver/pkg/storage/etcd3covering consistentGETs andGuaranteedUpdateconflict retries served from cache withConsistentGetFromCacheenabled and disabled.
Integration tests
- Integration tests in
test/integration/apiserver/correctnessverifyingGETconsistency andGuaranteedUpdateconflict handling acrossConsistentGetFromCachefeature gate toggles.
e2e tests
- Existing Kubernetes e2e and conformance suites exercise all
GET,CREATE,UPDATE, andPATCHoperations.
Graduation Criteria
Alpha
- Implement
ConsistentGetFromCachefeature gate to serve allGETs (includingGuaranteedUpdateconflictGETs) from the watch cache.
Beta
- Enable
ConsistentGetFromCacheby default.
GA
- Graduate
ConsistentGetFromCacheto GA and lock to default-on.
Upgrade / Downgrade Strategy
ConsistentGetFromCacheis a purely in-memory feature inkube-apiserverwith no persistent storage format changes.- Upgrading or downgrading
kube-apiserver, or toggling the feature gate, requires no data migration.
Version Skew Strategy
- Purely internal to each
kube-apiserverinstance. - Safe under HA mixed-version control planes (
v1.(N-1)andv1.N), as eachkube-apiserverindependently serves itsGETrequests from its own watch cache or etcd.
Production Readiness Review Questionnaire
Feature Enablement and Rollback
How can this feature be enabled / disabled in a live cluster?
- Feature gate (also fill in values in
kep.yaml)- Feature gate name:
ConsistentGetFromCache - Components depending on the feature gate:
kube-apiserver
- Feature gate name:
Does enabling the feature change any default behavior?
ConsistentGetFromCache: ConsistentGETrequests (resourceVersion="") and internalGETrequests triggered onGuaranteedUpdatetransaction conflicts are served from the watch cache. While the API semantics and consistency guarantees remain identical puts watch cache on critical path. This means that availability is dependent on the health of the watch cache and latency of those requests will be impacted by watch cache delay. However the incurrent latency is intentional to remove the costly fallback and allow APF to properly protect apiserver and etcd from overload.
Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?
Yes. ConsistentGetFromCache can be disabled by setting the feature gate to false and restarting kube-apiserver, which immediately reverts kube-apiserver to serving consistent GETs and TXN conflict retries directly from etcd.
What happens if we reenable the feature if it was previously rolled back?
kube-apiserver immediately resumes serving all GET requests from the watch cache.
Are there any tests for feature enablement/disablement?
ConsistentGetFromCache is purely in-memory feature, so enablement/disablement tests boil down to feature tests.
Rollout, Upgrade and Rollback Planning
How can a rollout or rollback fail? Can it impact already running workloads?
What specific metrics should inform a rollback?
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?
Monitoring Requirements
How can an operator determine if the feature is in use by workloads?
How can someone using this feature know that it is working for their instance?
- Events
- Event Reason:
- API .status
- Condition name:
- Other field:
- Other (treat as last resort)
- Details:
What are the reasonable SLOs (Service Level Objectives) for the enhancement?
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
- Metrics
- Metric name:
- [Optional] Aggregation method:
- Components exposing the metric:
- Other (treat as last resort)
- Details:
Are there any missing metrics that would be useful to have to improve observability of this feature?
Dependencies
Does this feature depend on any specific services running in the cluster?
Scalability
Will enabling / using this feature result in any new API calls?
No. Serving GETs from the watch cache reduces direct Range (GET) requests to etcd (especially during TXN conflicts), replacing full object fetches with lightweight watch progress checks when needed.
Will enabling / using this feature result in introducing new API types?
No.
Will enabling / using this feature result in any new calls to the cloud provider?
No.
Will enabling / using this feature result in increasing size or count of the existing API objects?
No.
Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?
No. It is designed to reduce GET, PATCH, and UPDATE latency and prevent degraded-state latency spikes under high write throughput.
Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?
No. It significantly reduces CPU and memory allocations in kube-apiserver and reduces read load on etcd.
Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?
No.
Troubleshooting
How does this feature react if the API server and/or etcd is unavailable?
What are other known failure modes?
What steps should be taken if SLOs are not being met to determine the problem?
Implementation History
- 2026-09-18: Initial write-throughput investigation and proposal filed in kubernetes/kubernetes#142223 .
- 2026-09-21: KEP-6404 drafted to serve all
GETs from cache (ConsistentGetFromCache) and document future serialization/deserialization allocation optimizations.
Drawbacks
- Puts the watch cache on the critical path for consistent
GETs (resourceVersion="") andGuaranteedUpdateconflict retries, making their latency and availability dependent on watch cache freshness. As noted in the PRR, this tradeoff is intentional to eliminate the expensive etcd decode fallback and allow APF to protectkube-apiserverand etcd from overload.
Alternatives
- Relying solely on
GetOnFailurein etcdTxnResponseforGuaranteedUpdateconflicts: While etcdTxnResponsecan return the conflicting key-value pair directly without an extra network round-trip,kube-apiserverstill has to deserialize and default those raw bytes into a fresh Go struct on every conflict. Serving the conflicting object from the watch cache returns the already-decoded object in memory, eliminating the per-conflict decode allocations.
Potential Future Extensions
Avoiding Whole-Object Decoding and Encoding
kube-apiserver has an ingrained bottleneck of operating on API objects as a whole rather than only on the relevant bytes that changed. While network transfer of large objects is rarely the bottleneck because wire payloads compress well, encoding and decoding whole 10–100KB objects consumes a massive share of CPU and memory allocations.
Consider a PATCH request to a Pod that only sets spec.nodeName or updates status: even though the change is only tens of bytes, kube-apiserver must serialize the entire 10–100KB Pod object. If kube-apiserver knew that only status was changing, it could serialize status alone and reuse the existing serialization of spec—reducing serialization cost from O(object size) to O(subtree that changed) at an appropriate subtree granularity. Similarly, on the read/watch path, kube-apiserver could reuse bytes already serialized in etcd instead of re-encoding unchanged subtrees.
Storage Versioning Metadata and Open Challenges
Most optimizations that reuse bytes stored in etcd share a common prerequisite: knowing that the bytes persisted in etcd are already up to date with the current kube-apiserver representation.
Although kube-apiserver already applies field defaulting on the write path (when decoding the incoming client request), it also applies defaulting on the read path to handle cluster upgrades: when a cluster upgrades from v1.(N-1) to v1.N and v1.N introduces a new field with a new default value, objects written before the upgrade do not have that default persisted in etcd. Because nothing in the etcd record indicates which Kubernetes version wrote the object, kube-apiserver currently must decode and default every object on read.
One candidate approach we explored was extending the etcd storage envelope metadata to record the Kubernetes minor version (k8sMinorVersion = max(storedVersion, currentVersion)) that wrote the object, using stored.k8sMinorVersion >= currentAPIServer.minorVersion to determine whether the stored serialization is up to date (analogous to how runtimes like Java fast-load bytecode produced by a compatible version without re-verification).
However, review identified several reasons why a Kubernetes minor-version stamp alone is insufficient to safely reuse or pass through undecoded bytes from etcd:
- Feature-gate-controlled defaulting: Default values can depend on whether specific feature gates are enabled on a given
kube-apiserverinstance, not just the binary’s minor version. - Unknown field stripping: Decoding into Go structs strips unknown fields, ensuring
kube-apiserveronly returns known fields to clients. Passing through or reusing undecoded bytes could preserve unknown fields and cause observable behavioral differences. - Storage-level decoration:
rest.Storagesupports read-timeDecoratorhooks (e.g.,store.Decorator = rest.defaultOnReadused in PVC and Service storage) that mutate objects on read. - Target version, format, and serialization options: Clients frequently request conversions across API versions, wire encodings (
JSONvs.Protobuf), or serialization options.
Alternative approaches—such as zero-allocation defaulting on the read path or subtree-level manipulation directly on Protobuf wire bytes—could reduce serialization/deserialization allocations from O(object size) toward O(changed fields) while respecting Kubernetes defaulting and unknown-field semantics, and will be explored in future proposals.