KEP-6283: Leveraging certificate APIs in the platform

Implementation History
ALPHA Implementable
Created 2026-08-18
Latest v1.38
Milestones
Alpha v1.38
Ownership
Owning SIG
SIG Auth
Primary Authors

KEP-6283: Leveraging new certificate APIs in the platform

Release Signoff Checklist

Items marked with (R) are required prior to targeting to a milestone / release.

  • (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
  • (R) KEP approvers have approved the KEP status as implementable
  • (R) Design details are appropriately documented
  • (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
    • e2e Tests for all Beta API Operations (endpoints)
    • (R) Ensure GA e2e tests meet requirements for Conformance Tests
    • (R) Minimum Two Week Window for GA e2e tests to prove flake free
  • (R) Graduation criteria is in place
  • (R) Production readiness review completed
  • (R) Production readiness review approved
  • “Implementation History” section is up-to-date for milestone
  • User-facing documentation has been created in kubernetes/website , for publication to kubernetes.io
  • Supporting documentation—e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes

Summary

Two KEPs that significantly improve Kubernetes ability to work with X.509 certificates went stable in Kube 1.37 - 3257-ClusterTrustBundles and 4317-PodCertificates. They lay out the APIs and implement the backing logic, but don’t directly make use of these. In this KEP, we will make use of the principles these APIs brought.

Motivation

We introduced mechanisms that are supposed to simplify people’s experience with X.509 certificates and bring their own signers but we don’t currently use these in places where they would make sense. This KEP does just that.

The KEP also aims to enhance the security of the Kubernetes ecosystem by providing built-in server TLS for service DNS names.

Goals

  • Direct use of ClusterTrustBundle references in API resources that mandate trust to a CA
  • Use ‘kubernetes.io/kube-apiserver-serving’ signer to mount kube-apiserver trust in pods’ ServiceAccount volumes
  • Introduce a new easy-to-use signer for server identies for Pods and Services

Non-Goals

  • remove the ‘kube-root-ca.crt’ ConfigMap from cluster namespaces
  • switch request-header-ca, client-ca from extension-apiserver-authentication to ClusterTrustBundles

Proposal

Allow users to use ClusterTrustBundle references in API resources that otherwise require a CA PEM inlined.

Add a serving signer that anyone can use in their workload to get a serving certificate trusted locally in the cluster.

Wire the ‘kubernetes.io/kube-apiserver-serving’ signer as the source of trust for kube-apiserver serving cert in the automounted ServiceAccount Pod volumes.

User Stories (Optional)

Story 1

Instead of pasting the PEM bundle in API objects that require it so that kube-apiserver can use it to trust a remote service, I would like to use a reference to a signer defined by a ClusterTrustBundle that is easier to maintain.

Story 2

I would like to have an easy-to-use way to get a cluster-local server identity for my pods so that I can securely serve content within the cluster.

Story 3

I would rather all pods used a single source for the ca.crt in their auto-injected ServiceAccount volumes. I like the positive outlook that we might one day be able to get rid of the kube-root-ca.crt ConfigMaps from all the namespaces and thus making life a little easier for our etcd databases.

Notes/Constraints/Caveats (Optional)

This enhancement removes the need for the existence of the kube-root-ca.crt ConfigMap as it uses a ClusterTrustBundle for its original purpose. However, we cannot just remove the ConfigMap as there might be people relying on it.

As for the serving signer, only a single signer exists for the whole cluster. If app developers need their own trust domain, they still need to create their own signer that would sign their PodCertificateRequests.

Risks and Mitigations

There are no known risks at this point, the features described are using stable Kubernetes features.

Design Details

Using ClusterTrustBundles directly as CA in API resources

This feature is controlled by the ClusterTrustBundleSelector feature gate that is off-by-default during alpha. It becomes on-by-default in beta and will be locked to be always enabled after GA.

There are a few cases of API resources where we allow injecting a PEM-formatted CA certificate, typically for the kube-apiserver to be able to use it while connecting to remote services. Namely, this occurs in these resources:

  • APIService - .spec.caBundle
  • {Mutating,Validating}WebhookConfig - .webhooks[].clientConfig.caBundle
  • CustomResourceDefinition - .spec.conversion.webhook.caBundle
  • kubectl config
    • these do not make sense to replace in normal kubeconfig, but there might be some value when for API servers’ static webhook configs
    • .clusters[].certificate-authority{,-data}
      • these should likely not be dynamically retrievable and should keep their current static form <– TODO: get agreement on that
    • .users[].client-certificate{,-data}
      • does not make sense to replace, we still need to provide the key somehow

For this purpose, we will create a new API structure that we put beside the API fields that expect CA PEM bundle strings. These fields will be mutually exclusive.

// ClusterTrustBundleSelector allows to configure trust by poiting to
// a selection of ClusterTrustBundles
type ClusterTrustBundleSelector struct {
    // name selects a single ClusterTrustBundle by name.
    // 
    // Mutually-exclusive with `signerName` and `labelSelector`.
    // +optional
    Name *string  `json:"name,omitempty"`
    
    // signerName selects all ClusterTrustBundles for a signer with
    // matching name.
    // The selection can be narrowed down by using `labelSelector`.
    // The contents of all selected ClusterTrustBundles will be
    // unified and deduplicated.
    //
    // Mutually-exclusive with `name`.
    // +optional
    SignerName *string `json:"signerName,omitempty"`
    
    // labelSelector allows to narrow down the selection of
    // ClusterTrustBundles for a signer with a given `signerName`.
    // If unset, interpreted as "match nothing". If set but empty,
    // interpreted as "match everything".
    //
    // Mutually-exclusive with `name`.
    // +optional
    LabelSelector *metav1.LabelSelector `json:"labelSelector,omitempty"`
}

The logic that handles the normalization of the resulting bundle MUST be shared with the relevant kubelet code that handles ClusterTrustBundles volume projection.

Platform-provided signer for serving certificates

This feature is controlled by the KubeServingCertificatesSigner feature gate that is off-by-default during alpha. It becomes on-by-default in beta and will be locked to be always enabled after GA.

A new certificate signer is introduced to handle PodCertificateRequests for serving certificates for Pods and Services - kubernetes.io/service-serving. The signer is represented by a controller in the Kubernetes Controller Manager that publishes a ClusterTrustBundle with the signer’s trust anchor, and another controller to issue certificates for incoming PodCertificateRequests that target this signer.

To be able to mint certificates on the service layer, the Kube Controller Manager receives new options:

  • --service-serving-key and --service-serving-cert that point to the signer’s key and certificate, respectively. These must be two different files.
  • --service-serving-trust-bundle is a file containing trust bundle in PEM format to be published inside of a ClusterTrustBundle for this signer.
  • --cluster-service-domain that will be appended for Pod/Service hostnames, e.g. my-pod.my-namespace.svc.cluster.local for service domain cluster.local

The signer expects at least one of the following parameters in PodCertificateRequest unverifiedUserAnnotations:

  • kubernetes.io/add-pod-fqdn
    • adds the Pod FQDN to the SAN DNS extension of the resulting certificate
    • the pod must set its spec.hostname and a headless Service whose name matches the Pod’s spec.subdomain must exist in the namespace and the Service’s selector must match the Pod
    • valid values: “true”
  • kubernetes.io/service-names
    • adds parameters from Kubernetes Service to the certificate
    • valid values: comma-separated list of Service names in the Pod’s namespace that match the Pod labels in the Service’s label selector

The certificates issued this way get KeyEncipherment | DigitalSignature key usages for RSA certificates and just DigitalSignature for others, and ServerAuth extended key usage.

Referencing Services for the kubernetes.io/service-serving signer

The kubernetes.io/service-names Service references require the following to be considered valid:

  • the referenced Service must exist in the Pod’s namespace
  • the spec.selector of the Service must match the Pod’s labels

If the above conditions are fulfilled, parameters of the Service are added into the resulting certificate.

For headless services the certificate gets SAN DNS names in the form of <svc-name>.<svc-namespace>.svc.<cluster-domain> and <svc-name>.<svc-namespace>.svc.

For cluster IP services the certificate gets SAN DNS names same as for headless services. The Service’s IP is NOT added to the certificate as SAN IP address to avoid situations where a service is removed and a different one that is created later reuses the original IP.

Handling failing certificate issuance preconditions

The signer controller cannot straigh out fail/deny a PodCertificateRequest when a Service that is either referenced directly or is expected to exist (see kubernetes.io/add-pod-fqdn) is missing in the cluster/does not have the expected shape. Failed/denied state is considered final for PodCertificateRequests and won’t be retried, which would cause the pod never to become running.

If a Service is not found (cached and live lookup) while handling a PodCertificateRequest, the signer controller issues an event for a pod that the PCR maps to and will consider the PCR handled for the time being.

The controller keeps an index that maps Services to PCRs. It watches Services and will trigger the signing loop for all unissued PCRs that map to a Service should a Service only be added/updated after a pod had been created.

On Pod IPs

It would seem natural to also be able to add Pod IPs in the certificate by using the same mechanisms as described above. Unfortunately, the Pod network is only set up once all the volumes are successfully mounted, specifically:

  1. Pod only gets .status.PodIPs before the volumes are handled if it is using host network
  2. Pod repeatedly waits for all volumes to be mounted properly
  3. control is handed over to the container runtime kubelet handler
  4. here the pod sandbox gets created and later Pod IPs are determined for the freshly created sandbox

As observed, Pod IPs are configured only after the volumes are set up. The codebase currently is not ready to handle situations where Pod network might be needed for some volumes, in fact this is likely a design choice.

Metrics

This feature exposes the following metrics:

  • service_serving_ca_controller_requests_total{result="issued|failed|error"}
    • a counter vector to track the total number of requests and their results
  • service_serving_ca_controller_signer_ttl_seconds
    • a gauge with unix timestamp marking the expiry of the signing certificate

Use the ‘kubernetes.io/kube-apiserver-serving’ signer in workloads for kube-apiserver trust

This feature is controlled by the KubeAPIServerWorkloadsTrust feature gate that is off-by-default during alpha and beta, but becomes on-by-default when the feature reaches GA.

Kubernetes automatically mounts a ServiceAccount token to pods via a projected volume. A part of that projected volume is a CA certificate to trust the kube-apiserver’s serving certificate.

This is done via admission that is using configmaps distributed by the kube-controller-manager for each namespace.

The ClusterTrustBundles KEP introduced a new signer for kube-apiserver serving trust - ‘kubernetes.io/kube-apiserver-serving’, along with a ClusterTrustBundle for it.

This KEP proposes to modify the admission plugin mentioned above to use the ClusterTrustBundle instead of the configmaps. In the future we may be able to remove the configmaps altogether but that is outside the scope of this proposal.

Test Plan

[x] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.

Prerequisite testing updates

-

Unit tests

  • k8s.io/apiserver/pkg/util/webhook: 2026-08-14 - 56.2%
  • k8s.io/kube-aggregator/pkg/apiserver/handler_proxy.go: 2026-08-14 - 78.7%
  • k8s.io/kubernetes/pkg/controller/certificates/rootcacertpublisher: 2026-08-14 - 69.1%

Integration tests

  • ClusterTrustBundleSelector
    • existing tests that check connectivity with CABundle will be extended to use the new fields
  • KubeServingCertificatesSigner
    • tests in test/integration/kubelet/podcertificatemanager_integration_test.go get extended with the new signer
  • KubeAPIServerWorkloadsTrust
    • currently there are no integration tests for ConfigMap behavior
      • extend just the unit test and e2e suite

e2e tests

  • ClusterTrustBundleSelector
    • existing tests that check connectivity with CABundle will be extended to use the new fields
  • KubeServingCertificatesSigner
    • tests in test/e2e/auth/projected_podcertificate.go get extended with the new signer
  • KubeAPIServerWorkloadsTrust
    • update tests that rely on ConfigMount mounts to make sure only ClusterTrustBundle projected mount is added on SA automounted pod volumes when the feature gate is enabled

Graduation Criteria

Alpha

  • Feature implemented behind a feature flag
  • Initial unit and e2e tests completed and enabled

Upgrade / Downgrade Strategy

The features expect the Cluster Trust Bundles and Pod Certificates feature gates to be enabled in all relevant components. Given the timeline of this KEP, by the time it goes stable the required feature gates will have been stable and on-by-default long enough so that we can rely on them being enabled even in the kubelet, which has the highest version skew.

Downgrading to versions without these features will require:

  • ClusterTrustBundleSelector - users need to remove all the new API fields from respective objects
  • KubeServingCertificatesSigner - users need to replace the new serving signer from their workloads’ manifests
  • KubeAPIServerWorkloadsTrust - no changes necessary.

Version Skew Strategy

As noted above, version skew should not be an issue for these features - they use stable APIs that should be available on the clusters at the time these features go GA.

Production Readiness Review Questionnaire

Feature Enablement and Rollback

How can this feature be enabled / disabled in a live cluster?
  • Feature gate (also fill in values in kep.yaml)
    • Feature gate name: ClusterTrustBundleSelector
    • Components depending on the feature gate: kube-apiserver
  • Feature gate (also fill in values in kep.yaml)
    • Feature gate name: KubeServingCertificatesSigner
    • Components depending on the feature gate: kube-controller-manager
  • Feature gate (also fill in values in kep.yaml)
    • Feature gate name: KubeAPIServerWorkloadsTrust
    • Components depending on the feature gate: kube-apiserver
Does enabling the feature change any default behavior?

Only KubeAPIServerWorkloadsTrust changes default behavior - instead of automounting a ConfigMap inside of pods that don’t explicitly opt-out, it mounts a ClusterTrustBundle.

Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?

Turning off KubeAPIServerWorkloadsTrust should have no effect on existing workloads.

Turning off KubeServingCertificatesSigner will stop rotations of any issued certificates and keys.

In case of ClusterTrustBundleSelector, users should replace the new API fields before turning the FG off by using the old caBundle fields. Failure to do so would result in failed kube-apiserver calls to the webhooks/API services that configured their trust using the new fields.

What happens if we reenable the feature if it was previously rolled back?

Everything should just start working again.

Are there any tests for feature enablement/disablement?

The behavior should all be unit tested with both feature gates on and off.

Rollout, Upgrade and Rollback Planning

How can a rollout or rollback fail? Can it impact already running workloads?
What specific metrics should inform a rollback?

KubeAPIServerWorkloadsTrust

  • sudden increase in failures to issue certificates, as tracked by service_serving_ca_controller_requests_total{result=~"failed|error"}, would likely mean there is an issue with the feature
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?

Monitoring Requirements

How can an operator determine if the feature is in use by workloads?
How can someone using this feature know that it is working for their instance?
  • Events
    • Event Reason:
  • API .status
    • Condition name:
    • Other field:
  • Other (treat as last resort)
    • Details:
What are the reasonable SLOs (Service Level Objectives) for the enhancement?
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
  • Metrics
    • Metric name:
    • [Optional] Aggregation method:
    • Components exposing the metric:
  • Other (treat as last resort)
    • Details:
Are there any missing metrics that would be useful to have to improve observability of this feature?

Dependencies

Does this feature depend on any specific services running in the cluster?

Scalability

Will enabling / using this feature result in any new API calls?
  • ClusterTrustBundleSelector

    • API call type: WATCH clustertrustbundles
    • estimated throughput: 1 per kube-apiserver
    • originating component: kube-apiserver
  • KubeServingCertificatesSigner

    • no new API calls
  • KubeAPIServerWorkloadsTrust

    • no new API calls
Will enabling / using this feature result in introducing new API types?

No, just new fields in existing API types for ClusterTrustBundleSelector.

Will enabling / using this feature result in any new calls to the cloud provider?

No.

Will enabling / using this feature result in increasing size or count of the existing API objects?
  • ClusterTrustBundleSelector

    • objects from Using ClusterTrustBundles directly as CA in API resources will include a new struct described in the section. While the resource grows, a typical object size should decrease since we no longer need to store the full string representation of a CA certificate.
      • no new objects
  • KubeServingCertificatesSigner

    • API type(s): ClusterTrustBundle
      • Estimated amount of new objects: 1
    • API type(s): PodCertificate
      • Estimated amount of new objects: 1 per certificate refresh period per pod using the signer
  • KubeAPIServerWorkloadsTrust

    • nothing new
    • brings a potential to remove a ConfigMap per namespace
Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?

No.

Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?

No.

Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?

No.

Troubleshooting

How does this feature react if the API server and/or etcd is unavailable?
What are other known failure modes?
What steps should be taken if SLOs are not being met to determine the problem?

Implementation History

Drawbacks

Alternatives

Infrastructure Needed (Optional)