KEP-6315: Weighted Load Balancing for Kube-Apiserver

Implementation History
ALPHA Provisional
Created 2026-09-02
Latest v1.38
Milestones
Alpha v1.38
Ownership
Participating SIGs
Primary Authors

KEP-6315: Weighted Load Balancing for Kube-Apiserver

Release Signoff Checklist

Items marked with (R) are required prior to targeting to a milestone / release.

  • (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
  • (R) KEP approvers have approved the KEP status as implementable
  • (R) Design details are appropriately documented
  • (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
    • e2e Tests for all Beta API Operations (endpoints)
    • (R) Ensure GA e2e tests meet requirements for Conformance Tests
    • (R) Minimum Two Week Window for GA e2e tests to prove flake free
  • (R) Graduation criteria is in place
  • (R) Production readiness review completed
  • (R) Production readiness review approved
  • “Implementation History” section is up-to-date for milestone
  • User-facing documentation has been created in kubernetes/website , for publication to kubernetes.io
  • Supporting documentation—e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes

Summary

In High-Availability (HA) Kubernetes clusters (e.g. 3-node control planes), traffic is distributed across kube-apiserver replicas through external or internal load balancers. Today, these load balancers operate with unweighted routing algorithms (such as round-robin or random distribution). However, control plane load across apiserver instances is inherently asymmetric. While on average traffic on nodes is distributed equally, leader elected controllers connect to usually only one kube-apiserver, establishing dozens of informers, persistent watch streams, and heavy write loops against single backend.

Because unweighted load balancers have no visibility into server load, they continue sending equal shares of external and cluster traffic to the already-saturated instance. This causes severe “hot node” degradation, latency spikes, API Priority and Fairness (APF) request rejections (HTTP 429), and potential node crashes, while other control plane replicas remain underutilized.

This KEP introduces Weighted Load Balancing support for kube-apiserver. By adopting the open Open Request Cost Aggregation (ORCA) standard, kube-apiserver communicates its real-time capacity and utilization to upstream load balancers via HTTP response headers. This standard is supported by layer 7 load balancers such as Envoy Proxy, Envoy Gateway, Istio, and cloud load balancers, allowing them to execute Weighted Round Robin (WRR) or Weighted Least Request (WLR) algorithms, automatically routing traffic to control plane nodes with available capacity and balancing utilization across the cluster.

Motivation

By enabling Weighted Round Robin (WRR) and Weighted Least Request (WLR) load balancing, the load balancer dynamically scales the routing weight of each kube-apiserver proportionally to its available headroom. When one apiserver is heavily loaded by active controllers, its weight is reduced, steering external traffic (from kubelets, CI/CD, user requests, and cluster add-ons) toward underutilized nodes and maintaining balanced control plane performance.

Goals

  • Adopt an open telemetry standard to expose real-time utilization metrics to load balancers.
  • Select a minimal set of utilization metrics that are validated to improve load distribution.
  • Improve load distribution for unary requests like POST, PATCH, DELETE, GET, LIST.
  • Provide a path to expand signals, if we see an opportunity to improve load distribution for other request types.

Non-Goals

  • Expose all possible resource and APF signals.
  • Implement client-side load balancing.

Proposal

We propose supporting Weighted Load Balancing for kube-apiserver by exposing real-time CPU and inflight request count via the Open Request Cost Aggregation (ORCA) telemetry standard using HTTP headers..

When enabled via the RequestCostAggregation feature gate and the --enable-endpoint-load-metrics flag, kube-apiserver will:

  1. Enable background goroutine to collect real time metrics, we expect around 100ms sampling rate should be a good balance of accuracy and overhead.
  2. Install a HTTP middleware that will attach standard ORCA headers (endpoint-load-metrics: TEXT cpu_utilization=0.3, mem_utilization=0.8, rps_fractional=10.0) to responses to requests that carry the endpoint-load-metrics-request header, see Securing Load Metrics .

User Stories

Story 1: Control plane behind a load balancer

As a cluster administrator running three kube-apiserver replicas that are reachable only through an L7 load balancer, I want the load balancer to weight traffic by apiserver utilization, without any API client being able to observe that utilization. Steps to configure weighted load balancing:

  1. enable load metrics on every kube-apiserver with the RequestCostAggregation feature gate and the --enable-endpoint-load-metrics flag,
  2. configure the load balancer to request load metrics by setting the endpoint-load-metrics-request header on every request it forwards to kube-apiserver, and
  3. configure the load balancer to drop endpoint-load-metrics and endpoint-load-metrics-bin from every response after consuming them, so they never reach clients.

With this in place, clients see no difference in the responses they receive.

Story 2: Clusters without an ORCA-aware load balancer

As an administrator of a cluster that does not run behind an ORCA-aware load balancer, for example a single apiserver or replicas behind an L4 load balancer, I don’t want kube-apiserver to reveal its resource utilization to clients. Load metrics require a dedicated flag that stays disabled by default, even after the feature graduates, so upgrading my cluster exposes nothing new, and clients that send the opt-in header receive no load metrics.

Notes/Constraints/Caveats

  • out-of-band reporting: For L7 load balancers that cannot inspect response or L4 load balancer, ORCA supports out-of-band reporting. Those usecases will be addressed at later Beta stage.
  • Header Size Overhead: The ORCA HTTP header adds a negligible payload (approx. 40–80 bytes) per HTTP response.

Risks and Mitigations

  • Exposing utilization to clients: load metrics reveal how busy each kube-apiserver instance is. This can help an attacker target the most loaded instance or time requests to periods of saturation, and can reveal the activity of other tenants.

    Mitigation: load metrics are never sent by default. The administrator has to enable them with the dedicated --enable-endpoint-load-metrics flag, and even then kube-apiserver attaches them only to responses to requests on which the load balancer negotiated them with the endpoint-load-metrics-request header. The load balancer removes them from responses before returning them to clients. See Securing Load Metrics . Clients that bypass the load balancer are discussed in Remaining exposure .

Design Details

Protocol Standard & Industry Adoption

The Open Request Cost Aggregation (ORCA) specification is an open, vendor-neutral standard created within the CNCF ecosystem. It was proposed in Envoy (envoyproxy/envoy#6614 ), and its load report schema and reporting service are maintained in cncf/xds . ORCA defines standard schemas and protocols for backend services to communicate utilization and request costs to load balancers.

ORCA is natively supported by:

  • Envoy Proxy: via built-in load balancing policies.
  • gRPC: via the ORCA OpenRcaService and out-of-band / in-band load report filters.
  • Istio / Service Meshes: via dynamic weighted routing based on backend capacity.

Securing Load Metrics

Utilization of a kube-apiserver instance is operational information that must not be available to every client. It can reveal the activity of other tenants and help an attacker target the most loaded instance or time requests to periods of saturation. Attaching load metrics to every response would make them readable by any client that can reach kube-apiserver, including anonymous clients, which can read endpoints such as /version and /readyz by default.

Load metrics are therefore attached to a response only when all of the following hold:

  1. The administrator enabled load metrics with the --enable-endpoint-load-metrics flag.
  2. The request carries the endpoint-load-metrics-request header, set by the load balancer.

The load balancer that sets the request header is responsible for removing load metrics from the response before returning it to its caller.

--enable-endpoint-load-metrics defaults to false, and keeps that default when the feature gate graduates. Enabling the feature gate by default in Beta, or locking it in GA, makes the capability available but never exposes load metrics on its own.

The flag is meant only for control planes fronted by a load balancer that consumes load metrics, in the same way --goaway-chance is meant only for kube-apiserver instances behind a load balancer.

Load balancer opt-in

ORCA does not standardize a request header for requesting load metrics. The ORCA design anticipates one without defining it: “A request header may indicate the custom metrics of interest, or alternatively the service will be preconfigured to send specific custom metrics.” gRPC A51 reports per-request metrics unconditionally, to “avoid the complexity of introducing some kind of negotiation for the client to tell the backend that it wants this data”. Unlike a typical gRPC backend, kube-apiserver serves untrusted clients, so it cannot report load unconditionally.

We propose the endpoint-load-metrics-request request header, named after the ORCA response headers endpoint-load-metrics and endpoint-load-metrics-bin. Its value selects the report format. Alpha supports only TEXT; requests with any other value receive no load metrics.

endpoint-load-metrics-request: TEXT

kube-apiserver does not forward the request header to backends it proxies requests to: aggregated API servers, and pods, services and nodes reached through proxy subresources. Before attaching its own report, it removes any ORCA response headers returned by those backends, so a proxied backend cannot inject or override the load report consumed by the load balancer.

Load balancer configuration

A load balancer that consumes load metrics from kube-apiserver must:

  • set endpoint-load-metrics-request on every request it forwards to kube-apiserver, and
  • remove endpoint-load-metrics and endpoint-load-metrics-bin from every response after consuming them, before returning it to its caller.

For example, with an Envoy route:

routes:
- match: { prefix: "/" }
  route: { cluster: kube-apiserver }
  request_headers_to_add:
  - header: { key: endpoint-load-metrics-request, value: TEXT }
    append_action: OVERWRITE_IF_EXISTS_OR_ADD
  response_headers_to_remove:
  - endpoint-load-metrics
  - endpoint-load-metrics-bin

Envoy Gateway removes both response headers by default: its BackendUtilization load balancer strips them unless keepResponseHeaders is set to true.

Remaining exposure

The request header signals intent; it does not authenticate the load balancer. With load metrics enabled, a client that reaches kube-apiserver directly, bypassing the load balancer, can send the header and read load metrics. Clients that go through the load balancer cannot, because the load balancer removes load metrics from every response.

Options:

  1. Document that load metrics may be enabled only when kube-apiserver is reachable exclusively through the load balancer, and accept the remaining risk.
  2. Honor the request header only on connections presenting a client certificate trusted for front-proxy authentication (--requestheader-client-ca-file, --requestheader-allowed-names), reusing the trust kube-apiserver already places in authenticating proxies. An L7 load balancer terminates TLS, so to preserve client certificate authentication it already has to act as such a proxy.
  3. Require the header value to carry a shared secret configured on both kube-apiserver and the load balancer.

Authorizing the request with RBAC does not help: behind an L7 load balancer the authenticated user is the end client, not the load balancer. «[/UNRESOLVED]»

Test Plan

[x] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.

Prerequisite testing updates
Unit tests
Integration tests
e2e tests

Graduation Criteria

Alpha

  • Feature gate RequestCostAggregation added (disabled by default).
  • Comprehensive unit and integration test coverage.

Beta

  • Feature gate RequestCostAggregation enabled by default.
  • Benchmark validation showing zero measurable throughput or latency regression in 5000-node scalability test suites.

GA

  • Feature gate RequestCostAggregation locked to true (unconditionally enabled).

Upgrade / Downgrade Strategy

To make use of the enhancement, an operator has to do two independent things: enable the RequestCostAggregation feature gate and the --enable-endpoint-load-metrics flag on kube-apiserver, and configure their L7 load balancer to request and consume ORCA load reports, remove them from responses, and use a weighted policy. Either step can be reverted independently, and reverting either one returns the cluster to unweighted balancing.

Version Skew Strategy

Telemetry will just be emitted from apiserver has the feature enabled.

Production Readiness Review Questionnaire

Feature Enablement and Rollback

How can this feature be enabled / disabled in a live cluster?
  • Feature gate (also fill in values in kep.yaml)
    • Feature gate name: RequestCostAggregation
    • Components depending on the feature gate: kube-apiserver
  • Other
    • Describe the mechanism: the --enable-endpoint-load-metrics flag on kube-apiserver, required in addition to the feature gate. See Securing Load Metrics .
    • Will enabling / disabling the feature require downtime of the control plane? No.
    • Will enabling / disabling the feature require downtime or reprovisioning of a node? No.

Enabling or disabling the gate or the flag requires a kube-apiserver restart.

Does enabling the feature change any default behavior?

No. Load metrics are attached to a response only when the administrator also sets --enable-endpoint-load-metrics and the request carries the endpoint-load-metrics-request header. No request handling, authorization, serialization, or response body changes in any way.

Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?

Yes. Disable the feature stops attaching load metrics to responses, causing loadbalancers to fallback to normal round robin algorithm.

What happens if we reenable the feature if it was previously rolled back?

Re-enabling should be exactly the same as first enabling as this is purely in memory feature.

Are there any tests for feature enablement/disablement?

The feature is purely in memory and keeps no persisted state, so enabling and disabling it is equivalent to running with the feature on or off, which is covered by the feature’s unit and integration tests.

Rollout, Upgrade and Rollback Planning

How can a rollout or rollback fail? Can it impact already running workloads?
What specific metrics should inform a rollback?
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?

Monitoring Requirements

How can an operator determine if the feature is in use by workloads?
How can someone using this feature know that it is working for their instance?
  • Events
    • Event Reason:
  • API .status
    • Condition name:
    • Other field:
  • Other (treat as last resort)
    • Details:
What are the reasonable SLOs (Service Level Objectives) for the enhancement?
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
  • Metrics
    • Metric name:
    • [Optional] Aggregation method:
    • Components exposing the metric:
  • Other (treat as last resort)
    • Details:
Are there any missing metrics that would be useful to have to improve observability of this feature?

Dependencies

Does this feature depend on any specific services running in the cluster?

Scalability

Will enabling / using this feature result in any new API calls?

No.

Will enabling / using this feature result in introducing new API types?

No.

Will enabling / using this feature result in any new calls to the cloud provider?

No.

Will enabling / using this feature result in increasing size or count of the existing API objects?

No. Nothing is persisted in etcd; the load report exists only on the wire.

Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?

No measurable increase is expected. The report is serialized once per sampling interval by a background goroutine and cached; the request path only performs an atomic load and sets one header, with no serialization and no allocation on the hot path.

This will be validated by benchmarks and by scalability tests before Beta.

Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?

No.

  • CPU: one background goroutine sampling process/cgroup CPU and memory counters at the proposed ~100ms interval; this is a small constant cost independent of QPS.
  • RAM: a single cached, pre-serialized report (tens of bytes) plus the sampler state.
  • Network: the header adds roughly 40-80 bytes per response when sent uncompressed. kube-apiserver serves HTTP/2, where HPACK indexes the header name once and indexes each distinct value, so the marginal cost for most responses is a few bytes; a new literal value is only sent once per sampling interval. At 10k QPS this is on the order of 0.1 MB/s, which is negligible relative to the volume of API response bodies and watch traffic at that scale.

The sampling interval is the main knob controlling both the sampler cost and the frequency of new HPACK literals, and the exact default will be validated during Alpha.

Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?

No. The feature adds a single goroutine and no new file descriptors, sockets, or connections.

Troubleshooting

How does this feature react if the API server and/or etcd is unavailable?
What are other known failure modes?
What steps should be taken if SLOs are not being met to determine the problem?

Implementation History

Drawbacks

Alternatives

Infrastructure Needed (Optional)