KEP-6313: Network-Isolated Pod Sandboxes
KEP-6313: Network-Isolated Pod Sandboxes
- Release Signoff Checklist
- Summary
- Motivation
- Proposal
- Design Details
- Production Readiness Review Questionnaire
- Implementation History
- Drawbacks
- Alternatives
Release Signoff Checklist
Items marked with (R) are required prior to targeting to a milestone / release.
- (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
- (R) KEP approvers have approved the KEP status as
implementable - (R) Design details are appropriately documented
- (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
- e2e Tests for all Beta API Operations (endpoints)
- (R) Ensure GA e2e tests meet requirements for Conformance Tests
- (R) Minimum Two Week Window for GA e2e tests to prove flake free
- (R) Graduation criteria is in place
- (R) all GA Endpoints must be hit by Conformance Tests within one minor version of promotion to GA
- (R) Production readiness review completed
- (R) Production readiness review approved
- “Implementation History” section is up-to-date for milestone
- User-facing documentation has been created in kubernetes/website , for publication to kubernetes.io
- Supporting documentation—e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes
Summary
This KEP adds network-isolated Pods: Pods that opt out of the default pod network. An isolated pod gets a network namespace with only a loopback interface and is not reachable from the cluster unless other network connectivity is added to it somehow.
Motivation
Kubernetes was built on the early assumption that every Pod is attached to a single flat cluster network and is routable from every other Pod. This assumption has always generated friction with workloads that need advanced routing, multi-tenancy or meshes. Those workloads get around the “default network” with out-of-band mechanisms: per-node runtime configuration, custom runtime classes, chained plugins, or annotations.
SIG Network has already tried to solve this inside the project. The Multi-Network effort (KEP-3698) concluded it was too complex: problems like network probes or DNS across multiple networks had no clear solution, portability depended on very specific (often private) plugin implementations, and the user stories were fragmented — some wanted multiple NICs, others VPC-like networks — each backed by a relatively small number of users. A follow-up proposal to simplify the low level API (KEP-4410 ) required offloading the network to another component that will make it incompatible with existing deployments and hard to roll out.
What has worked is composing on top of stable core primitives instead of extending the core network model:
- Kubernetes Network Drivers defined a new architecture, based on DRA, to provide a native mechanism to plumb additional network devices. DRANET is a project that demonstrated the successful use of this architecture to solve the challenges of the new AI/ML workloads and RDMA networks.
- KEP-4817 (GA) added a standard way to report network data in DRA. All drivers report their network devices with a common language, and integrators can use that data to build network integrations out of core (for example, their own EndpointSlice controllers), keeping core stable.
- NRI showed how network applications can integrate directly with the runtime. This removes the limitations of the CNI model, which composes plugins only through serial Unix-style chaining and file system based configuration that breaks the separation of concerns (see the service mesh CNI discussion ).
This friction is more critical these days. New workloads, like the agentic ones, require strict isolation, and their density and scale make the traditional out-of-band mechanisms obsolete. At the same time, DRA network drivers, VM sandboxes that provide their own networking (e.g. Kata/Firecracker/… with VFIO passthrough), disconnected jobs, … have been waiting for a solution for years.
One of the missing pieces is a core primitive to NOT attach the default pod network. Decomposing the Pod network attachment solves this problem for both maintainers and integrators: integrators get an isolated sandbox they can build on with the mechanisms above, and the project stops carrying the toil of workarounds and requirements that the core network model was never meant to express.
Goals
- Allow users to declare that a Pod must not be attached to the default pod network.
- Absorb the existing
hostNetworkfield into the new API so that existing clients and controllers keep working unchanged. - Fully define the semantics of isolated pods and fail predictably: reject configurations that cannot work without the pod network (network probes, host ports) at admission time instead of failing at runtime, and never silently run a pod that requested isolation attached to the pod network (e.g., on version-skewed nodes that do not support this feature).
- Keep isolated pods out of service discovery: Services, Endpoints, EndpointSlices and, by default, service environment variables.
- Define the contract between the kubelet and the container runtimes for isolated pods.
Non-Goals
- Designing a multi-network or secondary-interface attachment framework.
defaultNetwork: Noneis intentionally about the absence of automatic attachment; it provides a clean surface for delegation (e.g., DRA drivers), but how interfaces are attached afterwards and its lifecycle is out of scope. - Adding other networks to the API. Today
status.podIPs, Services, Endpoints and EndpointSlices, NetworkPolicy and cluster DNS only work with the default pod network, and this KEP does not change that (see Interaction with Core Networking APIs ). NewdefaultNetworkvalues for other networks, or changes to those APIs to support other networks, need a future KEP. - Isolating anything other than networking. Storage, IPC, PID, devices and host access are governed by existing fields.
- Deprecating or removing
hostNetwork. It is a GA v1 field and remains supported forever;defaultNetwork: Hostis its enum alias. - Changing how kube-proxy, NetworkPolicy or DNS work for non-isolated pods.
- Skipping service account token mounting for isolated pods. Tokens may still be consumed by non-network means.
Proposal
Add a new optional enum field defaultNetwork to PodSpec with the values
Pod, Host and None. When the field is unset, the API server defaults it
to Pod, or to Host when hostNetwork: true, so every existing pod gets
the value that describes its current behavior. None is the new behavior:
apiVersion: v1
kind: Pod
metadata:
name: network-isolated-example
spec:
defaultNetwork: None # new field; defaults to Pod (or Host when hostNetwork is true)
containers:
- name: worker
image: registry.k8s.io/e2e-test-images/agnhost:2.53
command: ["sleep", "3600"]
Semantics of defaultNetwork: None:
- The API server defaults
dnsPolicytoNone(without requiringdnsConfig.nameservers) andenableServiceLinkstofalse: isolated pods default to a configuration with no network dependencies. Users may override these defaults when an external integration (a DRA driver, VM networking) provides connectivity by other means, in line with the goal of offering a delegation contract to external integrations. Configurations that cannot work without a pod IP (network probes, host ports) are rejected by validation (see Validation ). - The kubelet uses a new CRI field to tell the runtime to create the sandbox network namespace with only the loopback interface and to skip the default-pod-network plumbing (e.g., runtimes that normally invoke CNI plugins to attach the sandbox to the default pod network would skip that step).
- The pod runs with empty
status.podIP/status.podIPsand becomesReadywithout waiting for an IP.status.hostIPis still reported. - The endpoints and endpointslice controllers never select the pod into Endpoints/EndpointSlices, even if a Service selector matches it.
- For isolated pods,
enableServiceLinksalso governs theKUBERNETES_*master-service variables, which are injected unconditionally for all other pods: with the defaultedfalseno service environment variables are injected, and an explicitenableServiceLinks: truerestores the standard injection for pods that obtain connectivity by other means. - The kubelet maps the pod hostname to
127.0.0.1and::1in the managed/etc/hosts, solocalhostand$(hostname)resolve inside the sandbox. - If the kubelet or the container runtime does not support isolated
sandboxes, the kubelet rejects the pod during node-level admission (after
scheduling, the pod is marked
Failedwith reasonPodFeatureUnsupported) instead of silently attaching it to the default pod network (fail closed). This is the existing Node Declared Features (KEP-5328) admission check; see Kubelet Behavior . - NetworkPolicy does not apply to these pods, as with
hostNetworkpods today: implementations enforce policy on pod network attachments and pod IPs, and an isolated pod has neither. Interfaces attached later by delegated mechanisms are outside the core network model and therefore outside NetworkPolicy scope. The full contract is defined in Interaction with Core Networking APIs .
defaultNetwork: Pod and defaultNetwork: Host do not change any behavior.
Pod is what pods do today: the pod gets its own network namespace and is
connected to the default pod network. Host is the same as hostNetwork: true.
User Stories
Story 1: DRA-provisioned accelerator fabric
As an operator of an AI training cluster, I allocate RDMA/accelerator NICs to pods through DRA drivers. I want the pod sandbox to start empty: no interfaces, routes or iptables rules from the default pod network plugin. The DRA driver is then the only owner of the pod connectivity. The default attachment is pure overhead for these pods: they communicate exclusively over the fabric interfaces, the default interface and its routes can conflict with the routing configured by the driver, and every sandbox pays IP allocation and network plumbing latency for a network it never uses. Today I have to coordinate teardown and priority with the CNI configuration on every node.
Story 2: Disconnected batch job
As a security engineer, I run jobs that process sensitive data mounted from
volumes, and I must prove to auditors that the workload has no network path
in or out. A deny-all NetworkPolicy depends on the network plugin enforcing
it, and the pod still gets a pod IP. With defaultNetwork: None there is
nothing to enforce: the only interface is loopback. The guarantee is about
provisioning, not a runtime invariant: the sandbox starts with only loopback
and defaultNetwork is immutable, so changing the pod’s connectivity afterwards
requires a privileged node-level actor (the workload itself cannot attach
interfaces without host-level privileges). Defending against privileged
actors is out of scope.
Story 3: VM sandboxes and device passthrough
As a confidential-computing platform owner using VM sandboxes, networking is provided inside the VM (e.g., via VFIO passthrough of a physical function). The node-level network plumbing is useless and potentially harmful. I want the runtime to skip it entirely, driven by a core API field rather than a runtime-specific annotation. Unlike Story 2, isolation is not the goal here: the platform itself provides the networking, and the empty sandbox is simply the correct starting state.
Story 4: Local-only multi-container pods
As a CI system author, I run pods whose containers only talk over localhost
(hermetic builds and tests). I want to guarantee that no external network
dependency can leak into the build, and I do not want these pods to consume
cluster IP space or appear in service discovery.
Story 5: Agentic orchestrators
As an operator of an agent orchestrator (e.g. Substrate), I run untrusted,
often machine-generated code in sandboxed pods at very high density and
churn. Each agent must start with zero network access by default; the
orchestrator then grants connectivity explicitly through its own channels
(a local proxy over localhost, or interfaces injected by a driver). I
cannot depend on NetworkPolicy enforcement for this guarantee, and at this
scale the per-pod network plumbing and IPAM are too much overhead: the pods never
use the default pod network, but they still consume IP space, endpoints
processing, and setup/teardown time on every sandbox.
Notes/Constraints/Caveats
Naming. This was extensively discussed during the initial reviews (kubernetes/kubernetes#141604 ):
- “Hermetic” was the original working name of the feature. It is not used in the API, the KEP title or this document: reviewers pointed out that “hermetic” implies storage, host access or compute isolation that this KEP does not provide.
- “CNI” never appears in the API. CNI is an implementation detail of Linux container runtimes; it does not exist in the Kubernetes API vocabulary.
defaultNetworkinstead ofnetworkModeorpodNetwork. Kubernetes has one network that every pod is connected to, the default pod network, and all the core networking APIs (Services, NetworkPolicy,status.podIPs, …) only work with that network. The field describes the pod’s default network: the default pod network (Pod), the host network (Host) or no network at all (None), sodefaultNetwork: Nonereads as “the pod has no default network”.Podinstead ofDefaultorClusterfor the current behavior: the pod gets its own network namespace and is connected to the default pod network. WithHostthe pod uses the host network namespace instead.defaultNetwork: Defaultis confusing, and “cluster network” is often understood to include the nodes.Nonefollows existing ecosystem semantics (docker run --network none) and the existingdnsPolicy: None. It describes the initial provisioning of the sandbox, not a permanent prohibition: interfaces can still be attached afterwards by delegated mechanisms (see Non-Goals ).
The pod still has a network namespace. The pod keeps loopback
(127.0.0.1, ::1) and the containers in the pod communicate over
localhost. This is why names like noNetwork were rejected: they are
technically inaccurate.
Kubernetes API access. An isolated pod cannot necessarily reach the
kube-apiserver (or anything else). The documentation in services-networking concepts
assumes
that every pod can reach the API and other pods. Those documents must be
updated to describe defaultNetwork: None as an explicit, opt-in exception.
Pods that do not set the field keep all the documented guarantees.
Downward API. status.podIP/status.podIPs resolve to empty values for
isolated pods. status.hostIP remains populated.
Secondary networks attached through CNI delegation. Some projects (for
example Multus
, and KubeVirt
through it) attach additional interfaces by
intercepting the CNI invocation that the runtime performs for the default pod
network. That invocation is an implementation detail of the runtime, and it
does not happen for defaultNetwork: None sandboxes, so today those projects
are not invoked for isolated pods. The semantics of None are “not attached
to the default pod network”, not “no secondary interfaces”: a runtime that
knows how to attach secondary interfaces without attaching the default pod
network is free to do so. Pods that do not set defaultNetwork: None are
not affected in any way, and existing integrations keep working unchanged
for them. For isolated pods, the paths that do not depend on the
default-pod-network CNI invocation are DRA network drivers (DRANET
) and
runtime hooks (NRI
, OCI hooks through CDI); aojea/network-device-plugin
demonstrates the pattern. Not breaking these downstream consumers and
providing them a compatibility path to isolated pods is a
Beta
graduation requirement.
Risks and Mitigations
Fail-open on version skew (security risk). The scheduler keeps isolated
pods away from unsupported nodes using Node Declared Features (KEP-5328)
,
but pods can bypass the scheduler (spec.nodeName, DaemonSets, static pods).
An old runtime or kubelet that does not understand the field could silently
attach an isolated pod to the default pod network. The kubelet fails closed
instead: if defaultNetwork is None and the node does not declare support
for isolated sandboxes (the gate is off, or the runtime does not report the
capability via CRI RuntimeFeatures), the kubelet rejects the pod at
admission with reason PodFeatureUnsupported. Kubelets that predate the field
cannot fail closed; for this reason the feature must not graduate to beta
(enabled by default) before every kubelet version within the supported skew
window knows the field. See
Version Skew Strategy
.
Ecosystem assumptions that every pod has an IP. Controllers, meshes and
operators that read status.podIP will see an empty value. For pods whose
sandbox is still being created this is already true today, but a Ready pod
with no IP is a new observable state, and components may mishandle it —
anywhere from logging errors to crashing (e.g., code that assumes
net.ParseIP(pod.Status.PodIP) returns non-nil for every Ready,
non-hostNetwork pod). A pod network implementation with that assumption
could be crashed by any user allowed to create pods. Mitigations: the
feature is opt-in behind a feature gate; before beta the e2e tests must
pass on the 4 most common pod network implementations (Calico,
OVN-Kubernetes, Kindnet and Cilium) with defaultNetwork: None pods in the
cluster; a conformance test checks this new state on every implementation;
and the new state is documented for integrators. Operators can additionally restrict who may set the field with
admission policy (e.g., ValidatingAdmissionPolicy). No Pod Security
Admission change is proposed: the field removes network access rather than
granting privilege. The endpoints controllers already ignore pods without
IPs.
API confusion with hostNetwork. Two fields describing the same thing is
not great, and the API change guidelines
warn about representing one
property with two fields, but hostNetwork is v1 and cannot go away.
PodSpec already carries two such pairs that are kept in sync in
SetDefaults_PodSpec: serviceAccount/serviceAccountName and
status.podIP/status.podIPs, where the older field is authoritative on a
mismatch for compatibility with older clients. defaultNetwork follows the
same rule (see Defaulting and hostNetwork mirroring
):
defaulting keeps both fields in sync, the only rejected combination is one
that old clients cannot produce, and a reader of either field gets the same
answer. Because defaulting runs before admission, Pod Security Admission
(which reads hostNetwork for the baseline host namespaces check) and
policy webhooks that inspect hostNetwork on Pods or workload templates see
the mirrored value and keep working without changes; defaultNetwork: Host
cannot be used to bypass them.
Scope creep towards multi-network. An enum makes it easy to propose
new values (bridge-like modes, named networks), and once the API names
the default pod network someone will want to name the other ones. Not in this
KEP: all the core networking APIs only work with the default pod network, and
that stays the same. Any new value or API change needs its own KEP and
cannot change the guarantees for pods connected to the default pod network.
Design Details
API Changes
New enum type and PodSpec field:
// PodDefaultNetwork describes the pod's default network.
// +enum
type PodDefaultNetwork string
const (
// PodDefaultNetworkPod gives the pod its own network namespace and
// attaches it to the default pod network (the network that Kubernetes
// connects every pod to unless the pod opts out). The container runtime
// performs its configured network plumbing and the pod is assigned pod
// IPs. This is the default and matches the historical behavior of
// Kubernetes.
PodDefaultNetworkPod PodDefaultNetwork = "Pod"
// PodDefaultNetworkHost runs the pod in the host's network namespace.
// It is equivalent to setting hostNetwork: true; the two fields are
// kept in sync by API defaulting.
PodDefaultNetworkHost PodDefaultNetwork = "Host"
// PodDefaultNetworkNone gives the pod its own network namespace
// containing only the loopback interface and does not attach it to the
// default pod network. The pod's podIPs are left unset. The pod is never
// selected into Services and, by default, receives no cluster DNS
// configuration or service environment variables.
PodDefaultNetworkNone PodDefaultNetwork = "None"
)
type PodSpec struct {
// ...
// defaultNetwork selects the pod's default network.
// "Pod" gives the pod its own network namespace attached to the default
// pod network, "Host" runs the pod in the host network namespace
// (equivalent to hostNetwork: true and kept in sync with it), and
// "None" gives the pod an isolated network namespace with only a
// loopback interface, not attached to the default pod network and with
// no automatic network plumbing.
// Defaults to "Pod", or to "Host" when hostNetwork is true; setting
// "Host" sets hostNetwork to true. "None" may not be combined with
// hostNetwork: true.
// When "None" is selected, dnsPolicy defaults to "None" and
// enableServiceLinks defaults to false (both may be overridden), and
// features that require networking (such as hostPorts and network-based
// probes and lifecycle handlers) are forbidden.
// This field is immutable.
// +featureGate=PodDefaultNetwork
// +optional
DefaultNetwork *PodDefaultNetwork `json:"defaultNetwork,omitempty" protobuf:"bytes,45,opt,name=defaultNetwork,casttype=PodDefaultNetwork"`
}
Why an enum and not a boolean. The first proposal was hermetic *bool
(see Alternatives
). It got a lot of pushback. A pod would
have two booleans, hostNetwork and hermetic, that cannot both be true,
so validation has to reject that combination, and adding a third option
later would need a third boolean. The API conventions
warn against this:
booleans “trend towards a small set of mutually exclusive options” and
rarely stay binary. With one enum a pod has exactly one value, the invalid
combination cannot be written, hostNetwork becomes the Host value, and a
new option later is just a new value.
The cost is keeping hostNetwork in sync, described in the next section.
hostNetwork is a GA v1 field and cannot be removed, so any design has to
deal with it. With the enum this happens in one place: old clients write
hostNetwork, new clients write defaultNetwork, and defaulting fills in
the other one.
Defaulting and hostNetwork mirroring
hostNetwork is a GA v1 field and cannot be removed. Mirroring an enum onto
a legacy boolean is the best we can do without breaking v1 compatibility. Defaulting
in pkg/apis/core/v1/defaults.go keeps both fields in sync, for Pods and
for every workload PodTemplateSpec:
func SetDefaults_PodSpec(obj *v1.PodSpec) {
// DefaultNetwork <-> HostNetwork mirroring.
if utilfeature.DefaultFeatureGate.Enabled(features.PodDefaultNetwork) {
switch {
case obj.DefaultNetwork == nil:
// Old clients: derive the mode from the legacy field.
if obj.HostNetwork {
obj.DefaultNetwork = ptr.To(v1.PodDefaultNetworkHost)
} else {
obj.DefaultNetwork = ptr.To(v1.PodDefaultNetworkPod)
}
case *obj.DefaultNetwork == v1.PodDefaultNetworkHost:
// New clients: keep the legacy field in sync for older
// controllers and clients that read hostNetwork directly.
obj.HostNetwork = true
case *obj.DefaultNetwork == v1.PodDefaultNetworkPod && obj.HostNetwork:
// "Pod" is what every stored object gets when the client did
// not set defaultNetwork, so "Pod" plus an explicit
// hostNetwork: true can only come from a client that does not
// know the new field (typically a PATCH on a workload
// template). hostNetwork: true keeps its historical meaning.
obj.DefaultNetwork = ptr.To(v1.PodDefaultNetworkHost)
}
}
if obj.DNSPolicy == "" {
if obj.DefaultNetwork != nil && *obj.DefaultNetwork == v1.PodDefaultNetworkNone {
obj.DNSPolicy = v1.DNSNone
} else {
obj.DNSPolicy = v1.DNSClusterFirst
}
}
if obj.DefaultNetwork != nil && *obj.DefaultNetwork == v1.PodDefaultNetworkNone &&
obj.EnableServiceLinks == nil {
obj.EnableServiceLinks = ptr.To(false)
}
// ... existing defaulting ...
}
Properties of this strategy:
- An old client submitting
hostNetwork: trueobservesdefaultNetwork: Hostafter defaulting; behavior is unchanged. - A new client submitting
defaultNetwork: HostproduceshostNetwork: truein the stored object, so an n-1 controller reading only the raw boolean makes the correct decision. - A client cannot send the conflict “
defaultNetwork: HostwithhostNetwork: false”.hostNetworkis a plain boolean, so the API server cannot tell an explicitfalseapart from an unset field. In both cases defaulting setshostNetworktotrueto matchdefaultNetwork: Host. hostNetwork: truealways meansHost. Once the gate is on, every stored Pod and workload template carriesdefaultNetwork: Podunless it asked for something else. An old client that patcheshostNetwork: trueinto such a template (kubectl patch, client-sidekubectl apply, or a controller built against an older client library) producesPodplushostNetwork: true; defaulting maps that toHost, so the client gets the pod it asked for instead of a validation error about a field it does not know.- The only conflict left is
defaultNetwork: Nonecombined with an explicithostNetwork: true, which validation rejects.Nonecan only have been set by a client that knows the field, so an old client hits this error only when it patcheshostNetwork: trueinto a template that a newer client already isolated; the error names both fields. - Setting one field from another has a known problem on updates: a client
that does not know the derived field cannot clear the source field with a
PATCH, because defaulting sets it back from the stored derived value. Both
hostNetworkanddefaultNetworkare immutable on Pods, so Pods are not affected. Workload templates are mutable, so an old client that patcheshostNetwork: falseinto a template that hasdefaultNetwork: Hoststored does not see its change. A full update (PUT) from the same client works, because the client drops the unknown field anddefaultNetworkis set again fromhostNetwork. The problem goes away once clients are updated. The API change guidelines describe a stricter pattern for paired fields (compare old and new objects in the registry strategy and derive whichever field did not change); it would have to be repeated in every registry that embeds aPodTemplateSpec, and the defaulting rule above covers the cases that matter for a boolean whose only conflicting value isNone. dnsPolicyandenableServiceLinksreceiveNone-specific defaults only when unset; explicit user values are preserved.- When the
PodDefaultNetworkfeature gate is disabled, the field is stripped from new objects (standarddropDisabledFieldshandling), no mirroring occurs, and behavior is exactly as today. Objects that already carry the field keep it, following the API compatibility rules for disabled gates; the cross-field validation below still runs on them, so a stored object is never left in a state that a later gate flip would reject.
Validation
Implemented in pkg/apis/core/validation/validation.go and invoked from
ValidatePodSpec (covering Pods and all workload templates):
// validatePodDefaultNetwork validates spec.defaultNetwork and its interactions
// with network-dependent fields.
func validatePodDefaultNetwork(spec *core.PodSpec, fldPath *field.Path) field.ErrorList {
allErrs := field.ErrorList{}
if spec.DefaultNetwork == nil {
return allErrs
}
netPath := fldPath.Child("defaultNetwork")
switch *spec.DefaultNetwork {
case core.PodDefaultNetworkPod:
if spec.HostNetwork {
// Defaulting maps this combination to "Host"; reaching this
// branch means an internal client bypassed defaulting.
allErrs = append(allErrs, field.Invalid(netPath, *spec.DefaultNetwork,
`must be "Host" when hostNetwork is true`))
}
case core.PodDefaultNetworkHost:
if !spec.HostNetwork {
// Defaulting keeps these in sync; reaching this branch means
// an internal client bypassed defaulting.
allErrs = append(allErrs, field.Invalid(netPath, *spec.DefaultNetwork,
"hostNetwork must be true when defaultNetwork is \"Host\""))
}
case core.PodDefaultNetworkNone:
allErrs = append(allErrs, validateIsolatedPodSpec(spec, fldPath)...)
default:
allErrs = append(allErrs, field.NotSupported(netPath, *spec.DefaultNetwork,
[]core.PodDefaultNetwork{core.PodDefaultNetworkPod, core.PodDefaultNetworkHost, core.PodDefaultNetworkNone}))
}
return allErrs
}
// validateIsolatedPodSpec enforces that a defaultNetwork "None" pod requests no
// feature that cannot work without a pod IP. dnsPolicy and enableServiceLinks
// are intentionally not validated: they are defaulted for "None" pods but
// remain overridable.
func validateIsolatedPodSpec(spec *core.PodSpec, fldPath *field.Path) field.ErrorList {
allErrs := field.ErrorList{}
if spec.HostNetwork {
allErrs = append(allErrs, field.Invalid(fldPath.Child("defaultNetwork"),
core.PodDefaultNetworkNone, "must not be \"None\" when hostNetwork is true"))
}
podshelper.VisitContainersWithPath(spec, fldPath, func(c *core.Container, cFldPath *field.Path) bool {
// The kubelet performs httpGet, tcpSocket and grpc probes from the
// host against the pod IP; without a pod IP there is nothing to
// probe. Only exec probes are permitted.
checkProbe := func(probe *core.Probe, probePath *field.Path) {
if probe == nil {
return
}
if probe.HTTPGet != nil {
allErrs = append(allErrs, field.Forbidden(probePath.Child("httpGet"),
`may not be set when defaultNetwork is "None"`))
}
if probe.TCPSocket != nil {
allErrs = append(allErrs, field.Forbidden(probePath.Child("tcpSocket"),
`may not be set when defaultNetwork is "None"`))
}
if probe.GRPC != nil {
allErrs = append(allErrs, field.Forbidden(probePath.Child("grpc"),
`may not be set when defaultNetwork is "None"`))
}
}
checkProbe(c.LivenessProbe, cFldPath.Child("livenessProbe"))
checkProbe(c.ReadinessProbe, cFldPath.Child("readinessProbe"))
checkProbe(c.StartupProbe, cFldPath.Child("startupProbe"))
if c.Lifecycle != nil {
checkHandler := func(h *core.LifecycleHandler, hPath *field.Path) {
if h == nil {
return
}
if h.HTTPGet != nil {
allErrs = append(allErrs, field.Forbidden(hPath.Child("httpGet"),
`may not be set when defaultNetwork is "None"`))
}
if h.TCPSocket != nil {
allErrs = append(allErrs, field.Forbidden(hPath.Child("tcpSocket"),
`may not be set when defaultNetwork is "None"`))
}
}
checkHandler(c.Lifecycle.PostStart, cFldPath.Child("lifecycle", "postStart"))
checkHandler(c.Lifecycle.PreStop, cFldPath.Child("lifecycle", "preStop"))
}
// hostPort mappings require an attached network interface.
portsPath := cFldPath.Child("ports")
for i, port := range c.Ports {
if port.HostPort > 0 {
allErrs = append(allErrs, field.Forbidden(portsPath.Index(i).Child("hostPort"),
`may not be set when defaultNetwork is "None"`))
}
}
return true
})
return allErrs
}
Additional validation adjustments:
dnsPolicyandenableServiceLinksare defaulted, not enforced. Isolated pods default to a configuration with no network dependencies (None,false), but external integrations that attach connectivity by other means (DRA drivers, VM networking) are part of the contract: users may override the defaults to match the connectivity those integrations provide.- Only
execprobes andexecorsleeplifecycle handlers pass validation.httpGet,tcpSocketandgrpcare rejected at admission because the kubelet performs them againststatus.podIP, and there is no pod IP. If a future mechanism lets the kubelet probe isolated pods, this validation can be relaxed. A reflection-based unit test walks the probe and handler types so that adding a new network-based probe without updating this validation fails CI. - Today,
dnsPolicy: Nonerequires the user to providednsConfig.nameservers. Pods withdefaultNetwork: Noneare an exception: they can have no nameservers at all. The kubelet then writes an empty/etc/resolv.conf, unless the user provides adnsConfig(for example, for a resolver listening on loopback). defaultNetworkcannot be changed after the pod is created, like the other fields that define the sandbox.ValidatePodStatusUpdaterejects a non-emptystatus.podIP/status.podIPswhenspec.defaultNetworkisNone, turning the empty-IP guarantee into an API invariant instead of a kubelet behavior. Status writes are normally node-owned and guarded only by RBAC and NodeRestriction; this narrow spec-consistency check also makes a fail-open node loud: its status updates are rejected instead of silently publishing an IP. The kubelet patches onlystatus, so the check applies to kubelets that do not know the field too. The consequence on such a node is that the pod’s status is never updated while its sandbox is running (it staysPendingin the API with the kubelet loggingFailed to update status for pod); the kubelet only reads sandbox IPs from aSANDBOX_READYsandbox, so once the pod is deleted and the sandbox stopped, the final status update succeeds and the pod is removed normally.
CRI Integration
A new typed field in PodSandboxConfig (k8s.io/cri-api) tells the runtime
whether to attach the sandbox network namespace to the default pod network.
CRI already expresses the host network through
NamespaceOption.network = NODE; the new field only describes sandboxes that
get their own network namespace (NamespaceOption.network = POD), so it has
no HOST value:
// PodSandboxDefaultNetwork selects whether a sandbox network namespace is
// attached to the default pod network.
// See https://kep.k8s.io/6313 for more details.
enum PodSandboxDefaultNetwork {
// DEFAULT_NETWORK_POD attaches the sandbox network namespace to the
// default pod network. The runtime performs its configured network
// plumbing (e.g., invokes CNI plugins on Linux). This is the default and
// matches historical behavior.
DEFAULT_NETWORK_POD = 0;
// DEFAULT_NETWORK_NONE creates the sandbox network namespace with only
// the loopback interface configured and does not attach it to the
// default pod network. The runtime MUST NOT perform its
// default-pod-network plumbing and MUST report an empty IP in
// PodSandboxStatus.
DEFAULT_NETWORK_NONE = 1;
}
message PodSandboxConfig {
// ... existing fields ...
// default_network selects whether the sandbox network namespace is
// attached to the default pod network. It only applies when
// NamespaceOption.network is POD; for NODE (host network) sandboxes
// there is no sandbox network namespace and runtimes MUST ignore this
// field. Kubelets that predate the field leave it unset
// (DEFAULT_NETWORK_POD).
// Feature gate: PodDefaultNetwork
// See https://kep.k8s.io/6313 for more details.
PodSandboxDefaultNetwork default_network = 10;
}
Proto3 cannot distinguish an unset enum from its zero value, so
DEFAULT_NETWORK_POD = 0 is what an older kubelet sends for every sandbox,
including host-network ones; this is why the field is scoped to sandboxes
with their own network namespace instead of mirroring the three Pod API
values. The values carry the DEFAULT_NETWORK_ prefix because proto enum
values share the package scope and POD is already a NamespaceMode value
in runtime.v1 (existing enums use the same convention: PROPAGATION_*,
SANDBOX_*, CONTAINER_*).
Capability discovery, so the kubelet can fail closed:
message RuntimeFeatures {
// ... existing fields ...
// default_network_none is set to true if the runtime supports creating
// sandboxes with PodSandboxDefaultNetwork DEFAULT_NETWORK_NONE.
bool default_network_none = 4;
}
Runtime obligations for DEFAULT_NETWORK_NONE sandboxes:
- create the netns according to
NamespaceOptionas usual; - bring up
lo(127.0.0.1/8,::1/128), the equivalent ofip link set lo up; - no default-pod-network plumbing on setup or teardown;
- report no IPs in
PodSandboxStatus.network.
A reference containerd implementation exists (aojea/containerd cniless
branch
). Upstream containerd and CRI-O implementations are a beta graduation
requirement.
Kubelet Behavior
- Admission is fail-closed, through Node Declared Features (KEP-5328)
.
The kubelet registers a
PodDefaultNetworkNonedeclared feature whoseDiscoverreturns true only when thePodDefaultNetworkgate is enabled and the runtime reportsdefault_network_noneinRuntimeFeatures, and whoseInferForSchedulingreturns true for pods withdefaultNetwork: None. This is the same shape as the existingUserNamespacesHostNetworkSupportfeature, which is gated on theuser_namespaces_host_networkruntime capability. The framework’s existing kubelet admit handler rejects a pod that requires a feature the node does not declare with reasonPodFeatureUnsupported(recorded as the pod’sstatus.reason, a warning event andkubelet_admission_rejections_total{reason="PodFeatureUnsupported"}), so adefaultNetwork: Nonepod is rejected when the gate is off or the runtime lacks the capability, with no kubelet code beyond the feature registration. Attaching an intentionally isolated pod to the default pod network is worse than not running it. generatePodSandboxConfigsetsdefault_network: DEFAULT_NETWORK_NONE.determinePodSandboxIPsreturns no IPs, andPodSandboxChangedmust not recreate the sandbox because the IP is missing.status.podIP/status.podIPsstay empty. Readiness is computed from container state only.status.hostIPis reported as usual.- The managed
/etc/hostsis still mounted, with the pod hostname (andhostname.subdomainif set) mapped to127.0.0.1and::1to avoid breaking applications that depend on that. - Service environment variables for isolated pods are governed entirely by
enableServiceLinks, including theKUBERNETES_*master-service variables that are injected unconditionally for other pods. With the defaultedfalseno service environment variables are injected and the kubelet skips the wait for the service informer to sync (serviceHasSynced) for these pods: one less control-plane dependency in the startup path. An explicitenableServiceLinks: truerestores the standard behavior, including the sync wait, for pods that obtain connectivity by other means. dnsPolicy: Nonewith nodnsConfigproduces an emptyresolv.conf. A user-supplieddnsConfigis honored: a resolver on loopback inside the pod, or nameservers reachable over driver-injected interfaces. An overriddendnsPolicy(for exampleClusterFirst) is honored exactly as for any other pod.
Control Plane Behavior
- No kube-controller-manager change is needed.
ShouldPodBeInEndpoints(ink8s.io/endpointslice/util, shared by the Endpoints and EndpointSlice controllers) already returnsfalsefor pods withoutstatus.podIPs, and isolated pods never have them (an API invariant, see Validation ), so they are never published even if a Service selector matches. A unit test makes this explicit fordefaultNetwork: Nonepods. - The
PodDefaultNetworkNonedeclared feature (see Kubelet Behavior ) is published innode.status.declaredFeatures, so the scheduler keeps isolated pods away from nodes without a supporting kubelet and runtime. This KEP only registers a new feature in that existing mechanism: it adds no scheduler code and no scheduler feature gate. TheNodeDeclaredFeaturesframework gate (GA and locked on since v1.37) is what enables the scheduler plugin that matches pod requirements againstnode.status.declaredFeaturesand the kubelet admit handler. The kubelet admission remains the final guarantee for pods that bypass the scheduler.
Interaction with Core Networking APIs
With defaultNetwork: None, “not connected to the default pod network”
becomes a supported state, so this section spells out how the core
networking APIs behave for those pods. All the core networking APIs work
only with the host network and the default pod network. The rules below
describe the current behavior; this KEP does not change any of them for
existing pods:
status.podIPsonly ever contains the pod’s IPs on the pod’s default network (pod-network IPs fordefaultNetwork: Pod, node IPs fordefaultNetwork: Host). Interfaces attached to isolated pods by delegated mechanisms are never reported there: publishing those addresses is the responsibility of the integration that attaches them (for example through the standardized DRA network device data of KEP-4817 ). Isolated pods therefore always have an emptystatus.podIPs, enforced by status validation (see Validation ).hostPortforwards from a host address tostatus.podIPsand nothing else. With no pod IPs there is nothing to forward to, which is why it is rejected by validation for isolated pods.- Service virtual IPs. Pods that are not attached to the default pod network are not required to be able to reach Service IPs, but they are allowed to: an implementation or integration may provide reachability by other means (a connect-time eBPF proxier, a driver-injected interface with the appropriate routes).
- Cluster DNS follows the same pattern for access: not required,
allowed. The
dnsPolicy: Nonedefault expresses “not required”; an explicit user override opting into cluster DNS is honored and works wherever the operator provides a path to the resolvers. The contents of cluster DNS are derived fromstatus.podIPs(through Endpoints and EndpointSlices), so isolated pods have no DNS records. - Service selection. Services select backends through
status.podIPs: the Endpoints and EndpointSlice controllers never publish isolated pods (ShouldPodBeInEndpoints), even when a selector matches. - NetworkPolicy
podSelector/namespaceSelectorselect traffic to and fromstatus.podIPsover the default pod network. Isolated pods have neither an attachment nor pod IPs, so NetworkPolicy does not apply to them, as with host-network pods today. - EndpointSlice API. In
addressType: IPv4/IPv6slices managed by the core controller, atargetRefto a Pod implies the addresses are that pod’sstatus.podIPs, reachable over the default pod network. Slices created by third-party controllers (such as the out-of-core integrations described in the Motivation) may carry addresses from other networks; that is existing behavior and unchanged by this KEP.
Changing any of these APIs to work with other networks, or adding
defaultNetwork values for other networks, is out of scope. It needs its
own KEP.
Test Plan
[x] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.
Prerequisite testing updates
None identified.
Unit tests
Planned coverage:
pkg/apis/core/validation:defaultNetworkcross-field validation (all enum values ×hostNetwork× probes × lifecycle handlers ×hostPort), the probe/handler completeness test, feature gate on/off, update immutability.pkg/apis/core/v1: defaulting ofdefaultNetwork,hostNetworkmirroring in both directions,dnsPolicy/enableServiceLinksdefaulting forNone, across Pod and all workload template kinds.pkg/api/pod:dropDisabledFieldsbehavior with the gate disabled.pkg/kubelet/kuberuntime: sandbox config generation, IP determination andPodSandboxChangedforDEFAULT_NETWORK_NONEsandboxes.staging/src/k8s.io/component-helpers/nodedeclaredfeatures/features:DiscoverandInferForSchedulingof thePodDefaultNetworkNonedeclared feature (gate on/off, runtime capability on/off).pkg/kubelet: managed hosts file content and environment variable construction for isolated pods.pkg/kubelet/status: readiness generation without pod IPs.staging/src/k8s.io/endpointslice/util:ShouldPodBeInEndpointsregression test with adefaultNetwork: Nonepod.pkg/apis/core/validation:2026-09-02-84.9%pkg/apis/core/v1:2026-09-02-81.2%pkg/kubelet/kuberuntime:2026-09-02-68.3%
Integration tests
test/integration/pods(TestPodDefaultNetwork*): create/update of Pods and of Deployment, StatefulSet, DaemonSet, Job and CronJob templates using every enum value; verification of defaulting and mirroring (including an old-client style PATCH ofhostNetwork: trueinto a template stored withdefaultNetwork: Pod, which must succeed and yieldHost), validation rejections (including the explicitdefaultNetwork: None+hostNetwork: trueconflict and the status update with a pod IP) and behavior with the feature gate disabled.test name : links to integration master and triage to be added once merged.
e2e tests
test/e2e/network/network_isolated.go(cluster e2e,Feature:PodDefaultNetwork, requires a runtime with support): pod runs and becomesReadywith emptystatus.podIP; onlyloexists in the sandbox; hostname resolves to loopback via/etc/hosts; no service environment variables (includingKUBERNETES_SERVICE_HOST) with the defaultedenableServiceLinks: false, injected when overridden totrue; downward APIstatus.podIPis empty; pod is excluded from EndpointSlices of a matching Service.test/e2e_node/network_isolated_test.go(node e2e): sandbox creation without network plumbing, fail-closed admission (PodFeatureUnsupported) against a runtime without support and with the kubelet gate disabled, kubelet restart idempotency (no sandbox recreation due to missing IP).Until upstream runtime releases exist, CI runs against kind with the patched containerd reference implementation; a periodic job will be added under sig-network.
test name : links to SIG Network testgrid and triage to be added once the job exists.
Graduation Criteria
Alpha
- Feature implemented behind the
PodDefaultNetworkfeature gate (default off) in kube-apiserver and kubelet. - new Pod field
defaultNetworkenum, defaulting/mirroring and validation complete. - CRI
PodSandboxConfig.default_networkandRuntimeFeatures.default_network_nonemerged incri-api. PodDefaultNetworkNoneregistered in Node Declared Features (KEP-5328) , which gives both the scheduler filter and the kubelet fail-closed admission.- Unit and integration tests listed above; initial e2e tests running against a reference runtime implementation.
Beta
- Promotion no earlier than three releases after the field is introduced, so that every kubelet version within the supported skew window (n-3) recognizes the field and fails closed, as required by the version skew policy.
- At least one released upstream runtime (containerd and/or CRI-O) supports
PodSandboxDefaultNetwork DEFAULT_NETWORK_NONEand reports the capability. - The e2e tests pass on clusters with the 4 most common pod network
implementations: Calico, OVN-Kubernetes, Kindnet and Cilium.
defaultNetwork: Nonepods work as described here and other pods are not affected: no crashes, no error log spam, NetworkPolicy keeps working. - Downstream consumers that attach secondary networks through CNI
delegation (Multus, and KubeVirt through it) are not broken: pods that do
not set
defaultNetwork: Nonekeep working unchanged with those projects installed, verified by e2e, and there is a documented and tested compatibility path (NRI or CDI hooks, or DRA) for them to attach interfaces todefaultNetwork: Nonepods. - Scheduling based on node declared features validated in heterogeneous clusters (mixed node versions and runtimes).
- Feedback gathered from DRA driver authors and early adopters.
- E2e tests running in Testgrid and linked in this KEP; no flakes for the gate-enabled jobs.
- Documentation updated, including the services-networking concepts pages that currently assume universal pod connectivity.
GA
- Container runtime support on containerd and CRI-O.
- Real-world usage of
defaultNetwork: Noneby at least two distinct classes of adopters (e.g., a DRA networking driver and a batch/security platform). - At least two releases after beta to allow user feedback and bug reports.
- Tests are promoted to conformance, so we guarantee that the default pod
network behavior does not change for
defaultNetwork: Pod/Hostpods, and the newdefaultNetwork: Noneguarantees the new behavior is implemented across all clusters. - All issues and gaps identified during beta resolved.
Deprecation
Not applicable
Upgrade / Downgrade Strategy
- Upgrade. Enabling the feature requires enabling the
PodDefaultNetworkgate on kube-apiserver and kubelets, plus a runtime version that reports the capability. Existing workloads require no changes:defaultNetworkdefaults to values consistent with their currenthostNetworksetting and behavior is identical. - Downgrade / gate disablement. New writes drop the field
(
dropDisabledFields); stored objects keep it, per the standard API compatibility rules. Kubelets with the gate disabled (including older kubelets that already know the field) refuse to admitdefaultNetwork: Nonepods instead of running them attached to the network. Isolated pods that are already running keep their sandboxes until they are deleted. - The mirroring rules guarantee that
hostNetworkremains a correct source of truth for any component that has not been upgraded.
Version Skew Strategy
- kube-apiserver (HA, mixed versions). During a rolling control-plane
upgrade, an n-1 apiserver does not know the field: it drops it on write,
or rejects the request when the client asks for strict field validation
(the
kubectldefault). This is the standard skew behavior for a new field; users should not rely on the field until all apiservers are upgraded and the gate is enabled everywhere. - Old clients. A client built before the field exists drops it on a full
update (PUT), so replacing a
defaultNetwork: Noneobject with such a client turns it back intodefaultNetwork: Pod; the pod then runs on the default pod network with thednsPolicy: None/enableServiceLinks: falsevalues it kept, which is visible. This is the general behavior of unknown fields with old clients; PATCH and server-side apply preserve the field. Old clients that only touchhostNetworkkeep working as described in Defaulting andhostNetworkmirroring . - kube-controller-manager (any version). The controller manager has no code for this feature. Isolated pods never have IPs, and pods without IPs are already excluded from Endpoints and EndpointSlices by existing controller logic.
- Old kubelet (n-1..n-3, field unknown). The most important skew: an old
kubelet drops the unknown field and would run the pod attached to the
default pod network (fail-open). All the mitigations work from alpha:
- old kubelets do not declare the node feature, so the scheduler never
places
defaultNetwork: Nonepods on them; - on upgraded nodes, the kubelet fail-closed admission covers the pods
that bypass the scheduler (
spec.nodeName, DaemonSets, static pods); - the gate is off by default, and beta promotion (default on) waits until all kubelet versions in the supported skew window know the field (see Graduation Criteria ). The residual risk is a pod that both bypasses the scheduler and targets a non-upgraded node while the cluster operator enabled the alpha gate; this is documented, and operators can cordon or upgrade those nodes.
- old kubelets do not declare the node feature, so the scheduler never
places
- New kubelet + old runtime. The runtime does not report
default_network_none, the node does not declare the feature, and the kubelet rejects the pod at admission (fail-closed,PodFeatureUnsupported). - hostNetwork consumers. Any component of any version reading
spec.hostNetworkobserves correct values thanks to mirroring.
Production Readiness Review Questionnaire
Feature Enablement and Rollback
How can this feature be enabled / disabled in a live cluster?
- Feature gate (also fill in values in
kep.yaml)- Feature gate name:
PodDefaultNetwork - Components depending on the feature gate: kube-apiserver, kubelet
- Feature gate name:
Does enabling the feature change any default behavior?
No. With the gate enabled, pods that do not set defaultNetwork receive a
defaulted value (Pod, or Host when hostNetwork: true) that exactly
describes their existing behavior. Only pods that explicitly set
defaultNetwork: None behave differently, and that value cannot be set today.
Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?
Yes. Disabling the gate on the apiserver causes the field to be dropped from
new objects; stored objects keep the field and remain valid. Kubelets with
the gate disabled refuse to admit defaultNetwork: None pods (fail closed);
running isolated pods keep their sandboxes until deleted. No
existing (non-isolated) workload is affected in any way by enabling or
disabling the gate.
What happens if we reenable the feature if it was previously rolled back?
The field becomes writable and enforceable again. Stored objects that
retained defaultNetwork values resume full semantics. There is no state to
reconcile beyond the field itself.
Are there any tests for feature enablement/disablement?
Yes. Unit tests cover dropDisabledFields and defaulting with the gate on
and off, including objects created with the field and then updated with the
gate disabled (field retention), and kubelet admission behavior for a
defaultNetwork: None pod with the gate disabled. Integration tests in
test/integration/pods exercise the same transitions against a real
apiserver.
Rollout, Upgrade and Rollback Planning
How can a rollout or rollback fail? Can it impact already running workloads?
Running workloads are not touched: the feature only affects pods explicitly
setting defaultNetwork: None, and sandboxes are never reconfigured in place.
Failure modes during rollout are limited to isolated pods themselves:
scheduling onto a node whose kubelet or runtime lacks support results in a
PodFeatureUnsupported admission rejection (visible, fail-closed) on
upgraded kubelets, or — on non-upgraded kubelets that drop the field — a pod
wired to the default pod network (see Version Skew). Mixed-version HA
control planes may intermittently drop the field on write until all
apiservers are upgraded.
What specific metrics should inform a rollback?
kubelet_admission_rejections_total{reason="PodFeatureUnsupported"}(unschedulable/failing isolated pods).- apiserver validation rejections of status updates that attempt to set
status.podIPondefaultNetwork: Nonepods (indicates a fail-open node that must be upgraded or cordoned). - Standard rollout signals:
kubelet_pod_start_duration_seconds, apiserver validation error rates on pod writes.
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?
Not yet; it will be tested before beta.
Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?
No.
Monitoring Requirements
How can an operator determine if the feature is in use by workloads?
Query pods with spec.defaultNetwork=None (e.g., via kube-state-metrics or a
field query). On nodes, PodFeatureUnsupported admission rejections (events
and kubelet_admission_rejections_total) indicate attempted use on
unsupported nodes.
How can someone using this feature know that it is working for their instance?
- API .status
- Other field: the pod is
Running/Readywith emptystatus.podIPandstatus.podIPs;kubectl execinto the pod shows only the loopback interface. Status updates attempting to report a pod IP for an isolated pod are rejected by validation and indicate a fail-open node.
- Other field: the pod is
- Events
- Event Reason:
PodFeatureUnsupportedwhen the node cannot honor the isolation request (fail closed); the message names the missingPodDefaultNetworkNonefeature.
- Event Reason:
What are the reasonable SLOs (Service Level Objectives) for the enhancement?
Isolated pod sandbox creation should be at least as fast and reliable as regular pods on the same node (it strictly removes the network plumbing step from the startup path). No regression in startup latency SLOs for non-isolated pods.
What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
- Metrics
- Metric name:
kubelet_pod_start_duration_seconds,kubelet_started_pods_errors_total,kubelet_admission_rejections_total{reason="PodFeatureUnsupported"}. - Components exposing the metric: kubelet
- Metric name:
Are there any missing metrics that would be useful to have to improve observability of this feature?
No. kubelet_admission_rejections_total already exists with a reason
label and PodFeatureUnsupported is already in its allowlist.
Dependencies
Does this feature depend on any specific services running in the cluster?
- Container runtime with CRI
PodSandboxDefaultNetwork DEFAULT_NETWORK_NONEsupport (containerd/CRI-O version supporting this KEP).
Scalability
Will enabling / using this feature result in any new API calls?
No.
Will enabling / using this feature result in introducing new API types?
No.
Will enabling / using this feature result in any new calls to the cloud provider?
No.
Will enabling / using this feature result in increasing size or count of the existing API objects?
Yes, marginally. PodSpec, and every PodTemplateSpec embedded in workload
objects, gains one optional enum field. With the gate enabled the field is
always populated by defaulting, so every Pod and template grows by a few
bytes (the field name and a short string value). No new objects are
created.
Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?
No. Isolated pod startup removes the network plumbing step entirely so it will improve pod startup.
Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?
No. Resource usage decreases using isolated pods (no CNI execution, no IPAM, no endpoints processing, no service environment variable construction by default).
Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?
No. Isolated pods consume fewer node resources than equivalent cluster-networked pods (no veth pairs, no IPs, …).
Troubleshooting
How does this feature react if the API server and/or etcd is unavailable?
By default, isolated pods do not depend on the service informer or cluster
DNS at startup: with enableServiceLinks: false no service environment
variables are constructed, so the kubelet can start them before services
have been listed. Otherwise standard kubelet behavior applies.
What are other known failure modes?
- Isolated pod stuck rejected on a node.
- Detection: pod
status.reasonand a warning event with reasonPodFeatureUnsupported;kubelet_admission_rejections_total{reason="PodFeatureUnsupported"}. - Mitigations: upgrade the node’s runtime/kubelet or reschedule the pod to a supporting node; this is the intended fail-closed behavior.
- Diagnostics: the event message lists the unsatisfied feature
(
PodDefaultNetworkNone);node.status.declaredFeaturesshows what the node supports; the kubelet log names the runtime and the missingRuntimeFeatures.default_network_nonecapability. - Testing: covered by node e2e against a runtime without support.
- Detection: pod
- Isolated pod attached to the default pod network (fail-open) on a
non-upgraded kubelet.
- Detection: the kubelet’s status updates are rejected by apiserver
validation (non-empty
status.podIPfor adefaultNetwork: Nonepod); the pod staysPendingin the API while running on the node, and the kubelet logsFailed to update status for pod. - Mitigations: cordon/upgrade the node; keep the gate off until all nodes are upgraded.
- Diagnostics: comparison of pod spec vs. status; node kubelet version.
- Testing: node declared features (from alpha) keep the scheduler away from these nodes; the fail-open path itself is not testable on the old kubelet.
- Detection: the kubelet’s status updates are rejected by apiserver
validation (non-empty
What steps should be taken if SLOs are not being met to determine the problem?
Inspect kubelet admission events/metrics for rejection reasons, verify runtime capability via the node’s CRI status, and check for version skew between apiserver, kubelet and runtime.
Implementation History
- 2026-08-26: Proposal socialized on the SIG Network mailing list and in a design doc , gathering initial feedback.
- 2026-08-26: POC published as kubernetes/kubernetes#141604 (API bool + validation/defaulting, CRI plumbing, kubelet behavior, endpoints exclusion, e2e/node tests, patched containerd reference implementation).
- 2026-09-02: KEP created as
provisional, adopting thenetworkModeenum design based on POC and design-doc review feedback. - 2026-09-09: Field renamed to
defaultNetworkwith valuesPod,HostandNonebased on KEP review feedback. - 2026-09-26: KEP marked
implementabletargeting alpha in v1.38. CRI field renamed todefault_network, scoped to sandboxes with their own network namespace; kube-controller-manager dropped from the feature gate; compatibility with CNI-delegation based secondary networks added as a beta criterion. - 2026-09-27: Fail-closed kubelet admission expressed as a Node Declared
Feature (
PodFeatureUnsupported, existing metric) instead of a bespoke check;hostNetwork: truetakes precedence over the defaultedPodvalue so old clients patching workload templates keep working; CRI enum values prefixed to avoid theNamespaceMode.PODclash.
Drawbacks
- For opted-in pods only, it breaks the long-standing documented model where every pod has an IP and can reach every other pod and the API server. The services-networking concepts documentation must be updated. Tooling that assumes every pod has an IP will see empty values (the same state as a sandbox that has not started yet).
- Container runtimes must implement and test a third sandbox networking mode.
Alternatives
A boolean field (hermetic: *bool). The first design, implemented in
the POC (kubernetes/kubernetes#141604
). It got a lot of pushback: a second
boolean that cannot be true at the same time as hostNetwork needs
validation for that combination and cannot grow to more options. The enum
says the same thing, the invalid combination cannot be written, and
hostNetwork becomes one of its values (see API Changes
).
Alternative field names. Considered and rejected:
| Candidate | Assessment |
|---|---|
hermetic: true | Implies isolation beyond networking (storage, host access, compute). Not used in the API, the KEP title or this document; see Notes/Constraints/Caveats . |
loopbackOnly: true | Too mechanical; describes the Linux namespace state rather than workload intent, and is unclear about DNS/service links/probes. |
isolated: true | Overloaded: confused with compute isolation (CPU pinning, NUMA), security sandboxes (gVisor/Kata) and NetworkPolicy namespace isolation. |
airgapped: true | Misleading: commonly refers to offline clusters/infrastructure, not individual sandboxes. |
disconnected: true | Sounds like a transient operational state (edge node offline), not an intentional configuration. |
noNetwork: true | Technically inaccurate: the pod does have a network namespace and a loopback interface. |
cniDisabled: true | Leaks a runtime implementation detail; CNI is not Kubernetes API vocabulary and does not apply to Windows or VM runtimes. |
podNetwork: true/false | A second boolean that cannot be true at the same time as hostNetwork; needs validation for that combination and cannot grow to more options. |
podNetwork: Default/None | Loses the hostNetwork absorption that motivates the enum, and pod.spec.podNetwork repeats the enclosing object name. |
networkMode: Default/Host/None | The previous version of this KEP. “mode” does not say what the pod is or is not connected to; defaultNetwork names the network, the same one all the core networking APIs work with. |
defaultNetwork: None avoids all of the above: it says which network the
pod is not connected to, None matches docker run --network none and
dnsPolicy: None, and the same field covers hostNetwork.
Pod annotation (e.g., io.kubernetes.cri.sandbox.hermetic). Annotations
are not versioned, not validated, and fail open: an unaware kubelet or
runtime silently ignores them, and the dependent fields (probes, DNS, service
links) cannot be validated.
Multi-network attachment API. The multi-network effort models an explicit list of networks that a pod attaches to; an empty list could mean “no networks”. That design is much larger and has unresolved conformance questions (Services, NetworkPolicy, probes and DNS across networks). This KEP only covers the absence of automatic attachment — a binary intent — and does not preclude any future multi-network work.
Deny-all NetworkPolicy. It depends on a plugin that enforces the policy. The pod still gets an IP and node CIDR space, still attaches interfaces, still publishes endpoints, and there is no guarantee about the sandbox contents.
RuntimeClass / per-node runtime configuration that skips CNI. It is not declarative per pod and it is invisible to the control plane: validation, endpoints and status handling remain wrong. It is not portable across runtimes and it requires dedicated node pools.