KEP-6315: Weighted Load Balancing for Kube-Apiserver
KEP-6315: Weighted Load Balancing for Kube-Apiserver
- Release Signoff Checklist
- Summary
- Motivation
- Proposal
- Design Details
- Production Readiness Review Questionnaire
- Implementation History
- Drawbacks
- Alternatives
- Infrastructure Needed (Optional)
Release Signoff Checklist
Items marked with (R) are required prior to targeting to a milestone / release.
- (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
- (R) KEP approvers have approved the KEP status as
implementable - (R) Design details are appropriately documented
- (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
- e2e Tests for all Beta API Operations (endpoints)
- (R) Ensure GA e2e tests meet requirements for Conformance Tests
- (R) Minimum Two Week Window for GA e2e tests to prove flake free
- (R) Graduation criteria is in place
- (R) all GA Endpoints must be hit by Conformance Tests within one minor version of promotion to GA
- (R) Production readiness review completed
- (R) Production readiness review approved
- “Implementation History” section is up-to-date for milestone
- User-facing documentation has been created in kubernetes/website , for publication to kubernetes.io
- Supporting documentation—e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes
Summary
In High-Availability (HA) Kubernetes clusters (e.g. 3-node control planes), traffic is distributed across kube-apiserver replicas through external or internal load balancers. Today, these load balancers operate with unweighted routing algorithms (such as round-robin or random distribution). However, control plane load across apiserver instances is inherently asymmetric. While on average traffic on nodes is distributed equally, leader elected controllers connect to usually only one kube-apiserver, establishing dozens of informers, persistent watch streams, and heavy write loops against single backend.
Because unweighted load balancers have no visibility into server load, they continue sending equal shares of external and cluster traffic to the already-saturated instance. This causes severe “hot node” degradation, latency spikes, API Priority and Fairness (APF) request rejections (HTTP 429), and potential node crashes, while other control plane replicas remain underutilized.
This KEP introduces Weighted Load Balancing support for kube-apiserver. By adopting the open Open Request Cost Aggregation (ORCA) standard, kube-apiserver communicates its real-time capacity and utilization to upstream load balancers via HTTP response headers. This standard is supported by layer 7 load balancers such as Envoy Proxy, Envoy Gateway, Istio, and cloud load balancers, allowing them to execute Weighted Round Robin (WRR) or Weighted Least Request (WLR) algorithms, automatically routing traffic to control plane nodes with available capacity and balancing utilization across the cluster.
Motivation
By enabling Weighted Round Robin (WRR) and Weighted Least Request (WLR) load balancing, the load balancer dynamically scales the routing weight of each kube-apiserver proportionally to its available headroom. When one apiserver is heavily loaded by active controllers, its weight is reduced, steering external traffic (from kubelets, CI/CD, user requests, and cluster add-ons) toward underutilized nodes and maintaining balanced control plane performance.
Goals
- Adopt an open telemetry standard to expose real-time utilization metrics to load balancers.
- Select a minimal set of utilization metrics that are validated to improve load distribution.
- Improve load distribution for unary requests like POST, PATCH, DELETE, GET, LIST.
- Provide a path to expand signals, if we see an opportunity to improve load distribution for other request types.
Non-Goals
- Expose all possible resource and APF signals.
- Implement client-side load balancing.
Proposal
We propose supporting Weighted Load Balancing for kube-apiserver by exposing real-time CPU and inflight request count via the Open Request Cost Aggregation (ORCA) telemetry standard using HTTP headers..
When enabled via the RequestCostAggregation feature gate and the --enable-endpoint-load-metrics flag, kube-apiserver will:
- Enable background goroutine to collect real time metrics, we expect around 100ms sampling rate should be a good balance of accuracy and overhead.
- Install a HTTP middleware that will attach standard ORCA headers (
endpoint-load-metrics: TEXT cpu_utilization=0.3, mem_utilization=0.8, rps_fractional=10.0) to responses to requests that carry theendpoint-load-metrics-requestheader, see Securing Load Metrics .
User Stories
Story 1: Control plane behind a load balancer
As a cluster administrator running three kube-apiserver replicas that are
reachable only through an L7 load balancer, I want the load balancer to weight
traffic by apiserver utilization, without any API client being able to observe
that utilization. Steps to configure weighted load balancing:
- enable load metrics on every
kube-apiserverwith theRequestCostAggregationfeature gate and the--enable-endpoint-load-metricsflag, - configure the load balancer to request load metrics by setting the
endpoint-load-metrics-requestheader on every request it forwards tokube-apiserver, and - configure the load balancer to drop
endpoint-load-metricsandendpoint-load-metrics-binfrom every response after consuming them, so they never reach clients.
With this in place, clients see no difference in the responses they receive.
Story 2: Clusters without an ORCA-aware load balancer
As an administrator of a cluster that does not run behind an ORCA-aware load
balancer, for example a single apiserver or replicas behind an L4 load
balancer, I don’t want kube-apiserver to reveal its resource utilization to
clients. Load metrics require a dedicated flag that stays disabled by default,
even after the feature graduates, so upgrading my cluster exposes nothing new,
and clients that send the opt-in header receive no load metrics.
Notes/Constraints/Caveats
- out-of-band reporting: For L7 load balancers that cannot inspect response or L4 load balancer, ORCA supports out-of-band reporting. Those usecases will be addressed at later Beta stage.
- Header Size Overhead: The ORCA HTTP header adds a negligible payload (approx. 40–80 bytes) per HTTP response.
Risks and Mitigations
Exposing utilization to clients: load metrics reveal how busy each
kube-apiserverinstance is. This can help an attacker target the most loaded instance or time requests to periods of saturation, and can reveal the activity of other tenants.Mitigation: load metrics are never sent by default. The administrator has to enable them with the dedicated
--enable-endpoint-load-metricsflag, and even thenkube-apiserverattaches them only to responses to requests on which the load balancer negotiated them with theendpoint-load-metrics-requestheader. The load balancer removes them from responses before returning them to clients. See Securing Load Metrics . Clients that bypass the load balancer are discussed in Remaining exposure .
Design Details
Protocol Standard & Industry Adoption
The Open Request Cost Aggregation (ORCA) specification is an open, vendor-neutral standard created within the CNCF ecosystem. It was proposed in Envoy (envoyproxy/envoy#6614 ), and its load report schema and reporting service are maintained in cncf/xds . ORCA defines standard schemas and protocols for backend services to communicate utilization and request costs to load balancers.
ORCA is natively supported by:
- Envoy Proxy: via built-in load balancing policies.
- gRPC: via the ORCA OpenRcaService and out-of-band / in-band load report filters.
- Istio / Service Meshes: via dynamic weighted routing based on backend capacity.
Securing Load Metrics
Utilization of a kube-apiserver instance is operational information that must
not be available to every client. It can reveal the activity of other tenants
and help an attacker target the most loaded instance or time requests to
periods of saturation. Attaching load metrics to every response would make
them readable by any client that can reach kube-apiserver, including
anonymous clients, which can read endpoints such as /version and /readyz by
default.
Load metrics are therefore attached to a response only when all of the following hold:
- The administrator enabled load metrics with the
--enable-endpoint-load-metricsflag. - The request carries the
endpoint-load-metrics-requestheader, set by the load balancer.
The load balancer that sets the request header is responsible for removing load metrics from the response before returning it to its caller.
--enable-endpoint-load-metrics defaults to false, and keeps that default
when the feature gate graduates. Enabling the feature gate by default in Beta,
or locking it in GA, makes the capability available but never exposes load
metrics on its own.
The flag is meant only for control planes fronted by a load balancer that
consumes load metrics, in the same way --goaway-chance is meant only for
kube-apiserver instances behind a load balancer.
Load balancer opt-in
ORCA does not standardize a request header for requesting load metrics. The
ORCA design
anticipates one without defining it: “A request header may indicate the custom
metrics of interest, or alternatively the service will be preconfigured to send
specific custom metrics.” gRPC
A51
reports per-request metrics unconditionally, to “avoid the complexity of
introducing some kind of negotiation for the client to tell the backend that it
wants this data”. Unlike a typical gRPC backend, kube-apiserver serves
untrusted clients, so it cannot report load unconditionally.
We propose the endpoint-load-metrics-request request header, named after the
ORCA response headers endpoint-load-metrics and endpoint-load-metrics-bin.
Its value selects the report format. Alpha supports only TEXT; requests with
any other value receive no load metrics.
endpoint-load-metrics-request: TEXT
kube-apiserver does not forward the request header to backends it proxies
requests to: aggregated API servers, and pods, services and nodes reached
through proxy subresources. Before attaching its own report, it removes any
ORCA response headers returned by those backends, so a proxied backend cannot
inject or override the load report consumed by the load balancer.
Load balancer configuration
A load balancer that consumes load metrics from kube-apiserver must:
- set
endpoint-load-metrics-requeston every request it forwards tokube-apiserver, and - remove
endpoint-load-metricsandendpoint-load-metrics-binfrom every response after consuming them, before returning it to its caller.
For example, with an Envoy route:
routes:
- match: { prefix: "/" }
route: { cluster: kube-apiserver }
request_headers_to_add:
- header: { key: endpoint-load-metrics-request, value: TEXT }
append_action: OVERWRITE_IF_EXISTS_OR_ADD
response_headers_to_remove:
- endpoint-load-metrics
- endpoint-load-metrics-bin
Envoy Gateway removes both response headers by default: its
BackendUtilization load balancer strips them unless keepResponseHeaders is
set to true.
Remaining exposure
The request header signals intent; it does not authenticate the load balancer.
With load metrics enabled, a client that reaches kube-apiserver directly,
bypassing the load balancer, can send the header and read load metrics. Clients
that go through the load balancer cannot, because the load balancer removes
load metrics from every response.
Options:
- Document that load metrics may be enabled only when
kube-apiserveris reachable exclusively through the load balancer, and accept the remaining risk. - Honor the request header only on connections presenting a client
certificate trusted for front-proxy authentication
(
--requestheader-client-ca-file,--requestheader-allowed-names), reusing the trustkube-apiserveralready places in authenticating proxies. An L7 load balancer terminates TLS, so to preserve client certificate authentication it already has to act as such a proxy. - Require the header value to carry a shared secret configured on both
kube-apiserverand the load balancer.
Authorizing the request with RBAC does not help: behind an L7 load balancer the authenticated user is the end client, not the load balancer. «[/UNRESOLVED]»
Test Plan
[x] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.
Prerequisite testing updates
Unit tests
Integration tests
e2e tests
Graduation Criteria
Alpha
- Feature gate
RequestCostAggregationadded (disabled by default). - Comprehensive unit and integration test coverage.
Beta
- Feature gate
RequestCostAggregationenabled by default. - Benchmark validation showing zero measurable throughput or latency regression in 5000-node scalability test suites.
GA
- Feature gate
RequestCostAggregationlocked to true (unconditionally enabled).
Upgrade / Downgrade Strategy
To make use of the enhancement, an operator has to do two independent things:
enable the RequestCostAggregation feature gate and the
--enable-endpoint-load-metrics flag on kube-apiserver, and configure their
L7 load balancer to request and consume ORCA load reports, remove them from
responses, and use a weighted policy. Either step can be reverted independently,
and reverting either one returns the cluster to unweighted balancing.
Version Skew Strategy
Telemetry will just be emitted from apiserver has the feature enabled.
Production Readiness Review Questionnaire
Feature Enablement and Rollback
How can this feature be enabled / disabled in a live cluster?
- Feature gate (also fill in values in
kep.yaml)- Feature gate name:
RequestCostAggregation - Components depending on the feature gate:
kube-apiserver
- Feature gate name:
- Other
- Describe the mechanism: the
--enable-endpoint-load-metricsflag onkube-apiserver, required in addition to the feature gate. See Securing Load Metrics . - Will enabling / disabling the feature require downtime of the control plane? No.
- Will enabling / disabling the feature require downtime or reprovisioning of a node? No.
- Describe the mechanism: the
Enabling or disabling the gate or the flag requires a kube-apiserver restart.
Does enabling the feature change any default behavior?
No. Load metrics are attached to a response only when the administrator also
sets --enable-endpoint-load-metrics and the request carries the
endpoint-load-metrics-request header. No request handling, authorization,
serialization, or response body changes in any way.
Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?
Yes. Disable the feature stops attaching load metrics to responses, causing loadbalancers to fallback to normal round robin algorithm.
What happens if we reenable the feature if it was previously rolled back?
Re-enabling should be exactly the same as first enabling as this is purely in memory feature.
Are there any tests for feature enablement/disablement?
The feature is purely in memory and keeps no persisted state, so enabling and disabling it is equivalent to running with the feature on or off, which is covered by the feature’s unit and integration tests.
Rollout, Upgrade and Rollback Planning
How can a rollout or rollback fail? Can it impact already running workloads?
What specific metrics should inform a rollback?
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?
Monitoring Requirements
How can an operator determine if the feature is in use by workloads?
How can someone using this feature know that it is working for their instance?
- Events
- Event Reason:
- API .status
- Condition name:
- Other field:
- Other (treat as last resort)
- Details:
What are the reasonable SLOs (Service Level Objectives) for the enhancement?
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
- Metrics
- Metric name:
- [Optional] Aggregation method:
- Components exposing the metric:
- Other (treat as last resort)
- Details:
Are there any missing metrics that would be useful to have to improve observability of this feature?
Dependencies
Does this feature depend on any specific services running in the cluster?
Scalability
Will enabling / using this feature result in any new API calls?
No.
Will enabling / using this feature result in introducing new API types?
No.
Will enabling / using this feature result in any new calls to the cloud provider?
No.
Will enabling / using this feature result in increasing size or count of the existing API objects?
No. Nothing is persisted in etcd; the load report exists only on the wire.
Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?
No measurable increase is expected. The report is serialized once per sampling interval by a background goroutine and cached; the request path only performs an atomic load and sets one header, with no serialization and no allocation on the hot path.
This will be validated by benchmarks and by scalability tests before Beta.
Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?
No.
- CPU: one background goroutine sampling process/cgroup CPU and memory counters at the proposed ~100ms interval; this is a small constant cost independent of QPS.
- RAM: a single cached, pre-serialized report (tens of bytes) plus the sampler state.
- Network: the header adds roughly 40-80 bytes per response when sent
uncompressed.
kube-apiserverserves HTTP/2, where HPACK indexes the header name once and indexes each distinct value, so the marginal cost for most responses is a few bytes; a new literal value is only sent once per sampling interval. At 10k QPS this is on the order of 0.1 MB/s, which is negligible relative to the volume of API response bodies and watch traffic at that scale.
The sampling interval is the main knob controlling both the sampler cost and the frequency of new HPACK literals, and the exact default will be validated during Alpha.
Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?
No. The feature adds a single goroutine and no new file descriptors, sockets, or connections.