KEP-6007: NUMA Utilization-Aware Topology Manager Policy Options

Implementation History
ALPHA Implementable
Created 2026-09-08
Updated 2026-09-22
Latest v1.38
Milestones
Alpha v1.38
Beta v1.39
Stable v1.41
Ownership
Owning SIG
SIG Node
Primary Authors

KEP-6007: NUMA Utilization-Aware Topology Manager Policy Options

Release Signoff Checklist

Items marked with (R) are required prior to targeting to a milestone / release.

  • (R) Enhancement issue in release milestone, which links to KEP dir in kubernetes/enhancements (not the initial KEP PR)
  • (R) KEP approvers have approved the KEP status as implementable
  • (R) Design details are appropriately documented
  • (R) Test plan is in place, giving consideration to SIG Architecture and SIG Testing input (including test refactors)
    • e2e Tests for all Beta API Operations (endpoints)
    • (R) Ensure GA e2e tests meet requirements for Conformance Tests
    • (R) Minimum Two Week Window for GA e2e tests to prove flake free
  • (R) Graduation criteria is in place
  • (R) Production readiness review completed
  • (R) Production readiness review approved
  • “Implementation History” section is up-to-date for milestone
  • User-facing documentation has been created in kubernetes/website , for publication to kubernetes.io
  • Supporting documentation, e.g., additional design documents, links to mailing list discussions/SIG meetings, relevant PRs/issues, release notes

Summary

Extend the Topology Manager with a numa-allocation-strategy policy option that controls NUMA node selection based on current allocation state. Valid values are none (default - preserves existing behavior), most-allocated (packing), and least-allocated (spreading).

These new policy options build on the Topology Manager policy options framework introduced in KEP-3545: Improved multi-numa alignment in Topology Manager . Like prefer-closest-numa-nodes from KEP-3545, the new options modify the behavior of existing Topology Manager policies (best-effort, restricted, single-numa-node) without introducing new policies.

Motivation

Modern multi-die CPU architectures expose multiple NUMA domains per socket (e.g., three compute dies per processor with Sub-NUMA Clustering enabled). For high-performance workloads such as packet-core network functions, each pod requires all of its allocated CPUs on a single NUMA node to maintain locality, while pod allocations must be balanced across the available NUMA nodes.

The Topology Manager currently selects NUMA nodes based purely on structural properties: affinity mask width (narrowest) or NUMA distance (closest, via KEP-3545). It has no awareness of how heavily each NUMA node is already allocated. The current allocation algorithm packs workloads onto the first NUMA node with sufficient available resources (lowest-ID tiebreak), resulting in unbalanced NUMA utilization (for example, four pods on NUMA 0 and two on NUMA 1) even when equivalent capacity exists elsewhere. This has been identified in kubernetes/kubernetes#125453 .

This creates two classes of problems:

  • NUMA imbalance on multi-die processors: On CPUs with 3+ NUMA domains per socket, the lowest-ID tiebreak concentrates workloads on one compute die while others sit idle, causing resource contention, turbo frequency imbalance, and degraded performance for latency-sensitive applications.
  • Fragmentation: Workloads spread across partially-used NUMA nodes, leaving no fully-free nodes available for large allocations that require full NUMA locality.

The closest existing option (distribute-cpus-across-numa) addresses a different use case: it spreads a single pod’s CPUs across multiple NUMA nodes, which is effectively the opposite of the desired behavior.

Operators today have no mechanism within the Topology Manager to influence NUMA node selection based on allocation state. This KEP addresses that gap.

Goals

  • Introduce a numa-allocation-strategy policy option for the Topology Manager that selects NUMA nodes based on current allocation state, with values none (default, no preference, preserves existing behavior), most-allocated (packing), and least-allocated (spreading).
  • Support per-resource weighting via numa-score-weights so operators can prioritize CPU, memory, or specific device types when computing utilization scores. Weights are specified as a comma-separated string of resource=weight pairs (e.g., "cpu=3,memory=1,nvidia.com/gpu=6"), where each weight is an integer in the range [0, 100]. Any resource not explicitly listed in the string defaults to a weight of 1, so naming a resource raises its influence rather than silencing the others.
  • Maintain existing topology guarantee semantics. The new options only influence NUMA node selection among equally valid candidates and do not change which hints are considered preferred.
  • Ensure no user-visible behavior change when the feature is disabled.

Non-Goals

  • Runtime resource utilization monitoring (e.g., actual CPU usage percent). This feature uses allocation state only.
  • Cross-node balancing or cluster-wide scheduling decisions. These remain the scheduler’s responsibility.
  • Device-specific utilization metrics (e.g., GPU memory usage). This feature uses simple device count for now.
  • Automatic weight tuning or workload profiling.

Proposal

Add a Score field to TopologyHint that captures per-NUMA-node allocation utilization. Hint providers (CPU manager, memory manager, device manager) each compute a score representing how heavily a NUMA node set is allocated. The Topology Manager then uses these scores, combined with new policy options, to prefer either the most-allocated or least-allocated NUMA nodes.

The design has three parts: score plumbing in TopologyHint and hint providers, allocation-aware policy options, and per-resource weight support via numa-score-weights. A score-aware merge optimization is planned for beta to improve performance on systems with many NUMA nodes.

Both options are configured per node, in topologyManagerPolicyOptions in the kubelet configuration, and apply to every pod admitted on that node. There is no per-pod or per-container opt-in and no pod API surface: a workload cannot request a different strategy than the one its node is configured with. This matches how the existing Topology Manager policy options work and is discussed further in Alternative 2 .

The unit the options are applied to is the one the Topology Manager already uses for hint merging, which topologyManagerScope (KEP-693 ) determines:

  • container scope (the default): hints are merged once per container, so the strategy is evaluated per container, against allocation state that already reflects the containers admitted before it.
  • pod scope: hints are merged once for the whole pod, so the strategy is evaluated once per pod and all containers share the resulting NUMA affinity.

This KEP introduces no new scoping concept and does not change which resources are aligned in either scope; it only changes which NUMA node is chosen among candidates that the existing logic already considers equivalent. The two scopes do, however, lead to different intra-pod placement, which is covered in Interaction with Topology Manager Scope .

User Stories

Story 1: Balanced NUMA Utilization on Multi-Die Processors

As a cluster operator running latency-sensitive network functions on processors with multiple compute dies (e.g., 3 or more NUMA domains per socket using sub-NUMA clustering), I want pod allocations to spread equally across all NUMA nodes so that no single compute die is overloaded while others sit idle.

Today, each pod requires all of its CPUs on a single NUMA node for performance locality, but the Topology Manager’s lowest-ID tiebreak concentrates pods onto NUMA node 0 until it is exhausted. On multi-die architectures this creates severe imbalance: one compute die runs near capacity while adjacent dies remain empty, leading to uneven turbo frequency behavior, thermal throttling, and wasted compute resources. The existing distribute-cpus-across-numa option addresses the opposite problem: it spreads a single pod’s CPUs across multiple NUMA nodes, which breaks the single-NUMA-node locality these workloads require.

With numa-allocation-strategy: "least-allocated", the Topology Manager would select the least-utilized NUMA node for each new pod, naturally balancing allocation across all compute dies while preserving per-pod NUMA locality.

See also: kubernetes/kubernetes#125453.

Story 2: Workload Consolidation for Power Efficiency

As a cluster operator running mixed batch and latency-sensitive workloads, I want batch jobs to pack onto already-utilized NUMA nodes so that empty nodes remain available for large latency-sensitive allocations requiring full NUMA locality. On multi-die processors, consolidating workloads onto fewer compute dies also allows idle dies to enter deeper power-saving states, reducing overall power consumption.

Story 3: Database Spreading

As a database administrator, I want database replicas to spread across under-utilized NUMA nodes to avoid thermal hotspots and improve sustained throughput across the machine.

Story 4: GPU Workload Prioritization

As an ML cluster operator, I want to weight GPU allocation higher than CPU when selecting NUMA nodes, so that GPU-heavy workloads prefer nodes where GPUs are already allocated, consolidating GPU usage even if CPU utilization is asymmetric. Without weights, CPU and memory pressure outvote the GPU signal and GPU-heavy pods fragment the free GPUs on the node.

Story 5: Excluding a Resource from Scoring

As an NFV operator, I want NUMA placement driven purely by SR-IOV VF availability, because VFs are the scarce non-fungible resource on my nodes: a NUMA node with no free VF cannot host the pod at all, while CPU and memory pressure is already handled by the scheduler. Giving CPU and memory a weight of 0 leaves the device provider as the only contributor to the score, so least-allocated steers pods toward the NUMA node with the most free VFs rather than the one that merely looks idle on CPU and memory.

Worked Examples: How Weights Influence Placement walks both of these scenarios through concrete allocation state and shows which NUMA node each weight string selects.

Notes/Constraints/Caveats

  • The policy options only affect NUMA node selection among candidates that are otherwise equivalent (same preferred status, same affinity width). They do not override the existing narrowest/closest tiebreakers. Score is only consulted when affinities are equal.
  • Score is based on allocation state (assigned CPUs, reserved memory bytes, allocated device count), not runtime utilization. This is a deliberate choice: allocation state is immediately available in the kubelet without additional monitoring infrastructure.
  • The most-allocated and least-allocated values for numa-allocation-strategy are mutually exclusive (the field accepts a single value).
  • numa-score-weights is coupled to numa-allocation-strategy. Weights only have an effect when numa-allocation-strategy is set to most-allocated or least-allocated. When it is none (or unset), scores are ignored entirely, so weights do nothing.
  • The effect of numa-allocation-strategy differs between topologyManagerScope: "pod" and topologyManagerScope: "container" (the default), because under container scope the strategy is applied once per container against allocation state that its predecessors have already updated. See Interaction with Topology Manager Scope .
  • numa-score-weights uses , and = as delimiters. Resource names containing = or , would break parsing. In practice, Kubernetes resource names follow DNS subdomain / slash / name conventions (e.g., nvidia.com/gpu, intel.com/sriov-nic) and do not use these characters, so this is not expected to be an issue.
  • These options compose with prefer-closest-numa-nodes from KEP-3545 : the comparison order is preferred flag, then narrowest/closest structural tiebreak, then score-based preference. This preserves structural guarantees while using score to break ties.

Risks and Mitigations

  • Score aggregation adds CPU overhead to hint comparison. Aggregation is O(n) where n is the number of providers (typically 2-4). Expected overhead is less than 1%.

  • Users misconfigure weights, causing unexpected behavior. The weight string is validated at kubelet startup. Malformed strings (e.g., missing = delimiter, non-numeric values) and weights outside the integer range [0, 100] are rejected with clear error messages. The kubelet will not start with an invalid configuration.

  • Bugs in the implementation lead to kubelet crash. Comprehensive unit and e2e testing will be used to mitigate this risk. If a crash does occur, disable the policy option and restart the kubelet. Pods already placed are not affected.

Design Details

Topology Hint Score Plumbing

Add a Score field to TopologyHint that captures allocation utilization:

type TopologyHint struct {
    NUMANodeAffinity bitmask.BitMask
    Preferred        bool
    Score            int64  // utilization score, 0=unscored, scored range [1,100]
}

Each hint provider computes a score per NUMA node set:

  • CPU manager: Score = max(1, assignedExclusiveCPUs * 100 / allocatableCPUs)
  • Memory manager: Score = max(1, assignedBytes * 100 / allocatableBytes)
  • Device manager: Score = max(1, allocatedDeviceCount * 100 / allocatableDeviceCount)

The Topology Manager aggregates these into a single score during hint merge. By default this is an equal-weight average, aggregatedScore = sum(scores) / count, which is the unweighted case of the general formula described in Per-Resource Weights .

Score is populated by all hint providers but does not influence hint selection on its own. It is only used when numa-allocation-strategy is set to most-allocated or least-allocated.

Allocation-Aware Policy Options

Add the numa-allocation-strategy policy option (with none, most-allocated, and least-allocated values) that makes Score a first-class selection criterion. When set to none (the default), Score is ignored and existing Narrowest/Closest behavior is preserved.

Updated comparison priority order:

  1. Preferred flag (topology constraint satisfaction, unchanged)
  2. Affinity mask comparison: Narrowest or Closest (structural tiebreak, unchanged)
  3. Score-based preference (if policy option is set, new)

This ensures that structural guarantees from prefer-closest-numa-nodes are preserved. Score only breaks ties among structurally equivalent candidates (same width, same distance).

Changes to hint comparison:

// In compareNumaAffinityMasks(current, candidate *TopologyHint) *TopologyHint:
// Step 1: Preferred always wins
if current.Preferred != candidate.Preferred {
    if candidate.Preferred { return candidate }
    return current
}

// Step 2: Structural comparison (Narrowest or Closest)
if best := compareStructural(current, candidate, opts); best != nil {
    return best  // Different width or distance, use structural
}

// Step 3: Score-based preference (only when structure is identical)
if opts.NUMAAllocationStrategy == "most-allocated" {
    if candidate.Score > current.Score { return candidate }  // higher = more packed
    if candidate.Score < current.Score { return current }
}
if opts.NUMAAllocationStrategy == "least-allocated" {
    if candidate.Score < current.Score { return candidate }  // lower = more empty
    if candidate.Score > current.Score { return current }
}

// Step 4: unscored (Score == 0) or equal scores: keep current, the same
// lowest-ID outcome the merge produces today
return current

When numa-allocation-strategy is "none" (or unset), Score is ignored and existing behavior is unchanged.

Interaction with Topology Manager Scope

The Topology Manager aligns resources either per container or per pod depending on topologyManagerScope, as defined in KEP-693 . Because numa-allocation-strategy consults allocation state rather than only structural properties, the two scopes lead to different intra-pod placement:

  • pod scope: hints are collected and merged once for the whole pod, so the strategy is applied a single time against one snapshot of allocation state. All containers in the pod share the resulting affinity, and the strategy has no intra-pod effect.

  • container scope (the default): hints are collected, merged, and allocated per container in sequence. Each container’s allocation is committed to the CPU, memory, and device managers before the next container’s hints are generated, so the scores seen by a later container already reflect the containers admitted before it. With least-allocated this means containers of the same pod tend to be placed on different NUMA nodes; with most-allocated they tend to be co-located.

The combination that warrants attention is therefore container scope with least-allocated, which is the only case where this feature makes containers of a pod less likely to share a NUMA node than they are today.

Intra-pod spreading in that combination is a consequence of the strategy rather than a violation of the topology guarantee. container scope promises per-container alignment only and has never promised a common NUMA node across containers. The co-location seen today is incidental rather than contractual: it falls out of the lowest-ID tiebreak and already breaks once the lowest-ID node is exhausted, at which point the next container spills onto another NUMA node. least-allocated does not remove a guarantee; it surfaces an existing non-guarantee earlier and more often.

Score quantization further limits how often this arises. Scores are integers in [1,100] computed as assigned * 100 / allocatable, so on a node with 128 CPUs per NUMA node a single-CPU container moves the score by 0.78, which truncates to 0. When the score does not change, the comparison falls through to the existing structural tiebreak and containers are placed as they are today. Spreading under container scope therefore only takes effect once a container is large enough relative to the NUMA node to move the score.

Preserving Current Placement Behavior

Operators who require containers of a pod to share a NUMA node have three options, none of which require changes to this design:

  • Leave numa-allocation-strategy as none. Placement is then unchanged.
  • Use most-allocated instead. Under container scope it strengthens co-location rather than weakening it: each container raises the score of the node it lands on, which attracts the next container to the same node. This is a stronger form of co-location than today’s tiebreak provides.
  • Set topologyManagerScope: "pod". This guarantees a common affinity across all containers, which is stricter than anything container scope offers today. It is not a drop-in substitution: pod scope admits a pod only if all of its containers can achieve a common alignment, so some pods that are admitted under container scope today will be rejected. That tradeoff is pre-existing KEP-693 behavior and is not introduced by this KEP. All example configurations in this KEP use pod scope.

Recovering pod-level co-location while remaining on container scope with least-allocated is explicitly not a goal. The policy interface is per-container by construction, and a rule that steered a container toward the NUMA node its siblings already occupy would reimplement pod scope inside container scope while contradicting what container scope means.

E2e tests will cover both scopes so that the difference is verified rather than incidental.

Score-Aware Preferred-First Merge Optimization (Beta)

This optimization is deferred to beta. In alpha, the existing Merge() logic is used with Score evaluated during comparison.

With Score driving selection, the merge can be made more efficient. Instead of exploring all O(N^K) hint permutations, a mergePreferred() method would explore only preferred hints sorted by Score so the best candidate is found first. On systems with many NUMA nodes (8+) and multiple providers, this reduces the search space significantly (e.g., from 49 permutations to 3 on a 3-NUMA, 2-provider system).

Per-Resource Weights

The numa-score-weights policy option is specified as a comma-separated string of resource=weight pairs in TopologyManagerPolicyOptions:

topologyManagerPolicyOptions:
  numa-score-weights: "cpu=3,memory=1,nvidia.com/gpu=6"

This keeps TopologyManagerPolicyOptions as map[string]string, consistent with all existing policy options. The string is parsed at kubelet startup into a map[string]int for internal use.

Weights are integers in the range [0, 100], the same form kube-scheduler’s NodeResourcesFit plugin uses for its per-resource weights.

The Topology Manager preserves resource names through the merge pipeline to enable weighted score aggregation. The providersHints structure is map[string][]TopologyHint where keys are resource names (“cpu”, “memory”, “nvidia.com/gpu”, etc.). During merge, resource names are threaded alongside hints so the aggregation function can look up weights[resourceName] for each provider’s score. This avoids duplicating resource names in every TopologyHint struct.

Scores are aggregated as a weighted average over the providers that reported a score for the candidate NUMA node set:

aggregatedScore = sum(weight[r] * score[r]) / sum(weight[r])

Because the denominator is the sum of the applicable weights, the aggregation is self-normalizing and the result stays in the same [1,100] range as the individual provider scores.

A provider the weight string does not name is given a weight of 1. Because 1 is also the smallest weight an operator can explicitly assign, an unnamed provider can never outweigh a named one: naming a resource raises its influence relative to everything else and never lowers it. When numa-score-weights is unset entirely every provider sits at 1, and the formula reduces to the equal-weight average sum(scores) / count described in Topology Hint Score Plumbing .

Since the baseline is 1, the only way to drop a resource from scoring altogether is to give it an explicit weight of 0.

The table below gives the effective weight of each provider for a pod requesting CPU, memory, GPU, and an SR-IOV NIC, where intel.com/sriov-nic is never named in the weight string:

ConfigurationCPUMemoryGPUSR-IOVEffect
(unset)1111Equal weighting, the same aggregation used when the option is absent
"nvidia.com/gpu=10"11101GPU counts ten times as much as each other provider
"cpu=3,memory=1,nvidia.com/gpu=6"3161GPU leads, CPU is intermediate, memory and the NIC sit at the baseline
"cpu=5,nvidia.com/gpu=5"5151CPU and GPU rank equally, memory and the NIC still contribute
"nvidia.com/gpu=10,cpu=0"01101GPU dominates and CPU is excluded outright
"cpu=100,memory=100,nvidia.com/gpu=100"1001001001The three named providers swamp the NIC without silencing it

Only the ratios between weights matter, so "cpu=30,memory=10,nvidia.com/gpu=60" and "cpu=3,memory=1,nvidia.com/gpu=6" behave identically; the smaller numbers are easier to read. A device type that appears on the node later contributes at weight 1 as soon as a pod requests it, which is the behavior we want for NUMA placement: a resource the operator has not ranked still affects locality, just less than the ones they have.

Worked Examples: How Weights Influence Placement

The examples below walk the weight strings through concrete allocation state and show which NUMA node each one selects. They are the basis for the user-facing documentation of this feature.

Example 1: prioritizing GPU consolidation (Story 4 ). A two-NUMA-node machine where each NUMA node has 64 exclusively allocatable CPUs, 128Gi of memory, and 4 GPUs, in the following allocation state:

NUMA nodeCPUs assignedMemory assignedGPUs assignedCPU scoreMemory scoreGPU score
numa048 / 6464Gi / 128Gi1 / 4755025
numa116 / 6432Gi / 128Gi3 / 4252575

A pod requests 8 exclusive CPUs, 16Gi of memory, and 1 GPU. Both NUMA nodes can satisfy it with a single-node affinity, so the two hints are structurally identical and the score decides. The operator runs most-allocated to consolidate GPU usage: filling numa1’s last GPU keeps a block of 3 free GPUs on numa0 for a later multi-GPU pod.

numa-score-weightsnuma0 aggregatenuma1 aggregateSelectedWhy
(unset)(75+50+25)/3 = 50(25+25+75)/3 = 41numa0CPU and memory pressure outvote the GPU signal, and the free GPU block is broken up
"nvidia.com/gpu=10"(75+50+250)/12 = 31(25+25+750)/12 = 66numa1GPU dominates, so the nearly-full GPU node wins
"cpu=3,memory=1,nvidia.com/gpu=6"(225+50+150)/10 = 42(75+25+450)/10 = 55numa1GPU still leads, CPU moderates it
"cpu=5,nvidia.com/gpu=5"(375+50+125)/11 = 50(125+25+375)/11 = 47numa0CPU ranked equal to GPU hands the decision back to CPU utilization
"nvidia.com/gpu=10,cpu=0"(0+50+250)/11 = 27(0+25+750)/11 = 70numa1CPU excluded outright, GPU and memory decide
"cpu=100,memory=100,nvidia.com/gpu=100"5041numa0Uniform weights are equivalent to leaving the option unset

Two properties are worth calling out. A weight expresses influence relative to the other providers, so ranking CPU as highly as GPU (row 4) reverses the decision even though the GPU weight did not change. And because the aggregate is self-normalizing, scaling every weight by the same factor (row 6) changes nothing.

The strategy and the weights answer different questions: the weights decide which resource’s utilization matters, the strategy decides which direction to move along it. With "nvidia.com/gpu=10" and least-allocated instead, the same state selects numa0 (31 < 66), the node with the most free GPUs.

Example 2: excluding a resource from scoring (Story 5 ). An NFV node with four NUMA nodes, each with 32 exclusively allocatable CPUs, 128Gi of memory, and 8 SR-IOV VFs. The operator runs least-allocated and cares only about VF availability, because a NUMA node with no free VF cannot host the pod at all while CPU and memory pressure is already handled by the scheduler.

NUMA nodeCPU scoreMemory scoreVF scoreFree VFsUnweighted aggregate"cpu=0,memory=0"
numa025257524175
numa175752565825
numa250505045050
numa350505045050

With weights unset, least-allocated selects numa0, the node with the fewest free VFs, because its low CPU and memory utilization pulls the average down. Repeated placements exhaust numa0’s VFs while numa1 keeps 6 idle. Zeroing CPU and memory leaves intel.com/sriov-nic as the only applicable provider at the default weight of 1, so the aggregate equals the VF score and least-allocated selects numa1.

Zeroing is the sharpest instrument the option offers and it has a corresponding edge: a pod that requests none of the providers left with a non-zero weight has a total applicable weight of 0, and its placement falls back to the existing Narrowest/Closest tiebreak.

If the total applicable weight comes out as 0 the weighted average is undefined. In practice this means no provider reported a score, which is the case for a pod that requests none of the resources the score-aware providers cover. The Topology Manager then compares hints using the existing Narrowest/Closest logic alone, the behavior it already has when numa-allocation-strategy is "none".

Provider score weights are validated when the kubelet configuration is parsed:

  • All weights must be integers in the range [0, 100]. Negative, fractional, and out-of-range values are rejected with an error. A weight of 0 is accepted and excludes the resource from scoring.
  • Resource names are not validated against the providers present on the node. The applicable provider set is a property of the pod being admitted rather than of the node, so a name that matches nothing on the node is not distinguishable from a name that simply is not requested by a given pod. A weight string naming a resource no workload requests is accepted; that resource never contributes a score.

Kubelet Configuration

Both new policy options are string values in the existing TopologyManagerPolicyOptions (map[string]string). No type change is needed.

KeyValueDescription
numa-allocation-strategy"none", "most-allocated", "least-allocated"Controls NUMA node selection based on allocation state
numa-score-weights"resource=weight,..." (e.g., "cpu=3,nvidia.com/gpu=6")Comma-separated resource=weight pairs (integers in [0, 100]) for weighted score aggregation

Example Configurations

No preference (default, preserves existing behavior):

apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
topologyManagerPolicy: "single-numa-node"
topologyManagerScope: "pod"
topologyManagerPolicyOptions:
  numa-allocation-strategy: "none"

Packing (consolidate workloads):

apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
topologyManagerPolicy: "single-numa-node"
topologyManagerScope: "pod"
topologyManagerPolicyOptions:
  numa-allocation-strategy: "most-allocated"

Spreading (balance load):

apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
topologyManagerPolicy: "restricted"
topologyManagerScope: "pod"
topologyManagerPolicyOptions:
  numa-allocation-strategy: "least-allocated"

GPU-weighted packing:

apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
topologyManagerPolicy: "single-numa-node"
topologyManagerScope: "pod"
topologyManagerPolicyOptions:
  numa-allocation-strategy: "most-allocated"
  numa-score-weights: "nvidia.com/gpu=6,cpu=3,memory=1"

Feature Gate

Name: TopologyManagerPolicyAlphaOptions (existing gate from KEP-3545 , reused for new alpha-stage options)

Behavior when disabled:

  • numa-allocation-strategy is ignored if specified.
  • numa-score-weights is ignored if specified.
  • Score field exists in TopologyHint but is not used in hint comparison.
  • Existing Narrowest/Closest behavior is preserved.

Test Plan

[x] I/we understand the owners of the involved components may require updates to existing tests to make this code solid enough prior to committing the changes necessary to implement this enhancement.

Prerequisite testing updates

None. Existing Topology Manager test infrastructure is sufficient.

Unit tests
  • k8s.io/kubernetes/pkg/kubelet/cm/topologymanager: 2026-09-16 - 90.8

as reported by go test -cover ./pkg/kubelet/cm/topologymanager/. This is the current baseline; we’ll re-measure once the implementation lands and ensure the new scoring and policy option code maintains or improves it.

Unit tests will cover:

  • Score calculation for each provider (CPU/Memory/Device)
  • Score aggregation with various weight combinations (equal-weight, CPU-weighted, GPU-weighted)
  • Weight validation: range [0, 100], rejection of negative, fractional, and out-of-range values
  • Weight normalization: verify "cpu=30,memory=10,nvidia.com/gpu=60" yields the same result as "cpu=3,memory=1,nvidia.com/gpu=6"
  • Default weight resolution: verify all providers are weighted equally when numa-score-weights is unset, and that unlisted providers default to weight 1 when it is set
  • Explicit exclusion: verify weight=0 excludes a resource from scoring
  • Policy comparison with most/least-allocated preferences
  • Edge cases: no scores, equal scores, and a container that requests none of the score-reporting providers (no applicable weight, expect fallback to Narrowest/Closest logic)
  • Mixed named and unnamed providers: a container requesting a resource the weight string does not name is still scored over that provider at weight 1, and the self-normalizing denominator keeps the aggregate in the [1,100] range
  • Validation of numa-allocation-strategy values (none, most-allocated, least-allocated)
  • Feature gate on/off behavior
e2e tests

E2e tests will cover:

  • Pod admission with packing policy option (verify NUMA node selection prefers most-allocated nodes)
  • Pod admission with spreading policy option (verify NUMA node selection prefers least-allocated nodes)
  • Default behavior preserved when numa-allocation-strategy is "none" or unset
  • Both topologyManagerScope: "pod" and topologyManagerScope: "container", verifying that a multi-container pod shares a NUMA node under pod scope and that container scope applies the strategy per container as described in Interaction with Topology Manager Scope

Graduation Criteria

Alpha

  • Feature implemented behind TopologyManagerPolicyAlphaOptions feature gate
  • Score plumbing in TopologyHint and hint providers
  • numa-allocation-strategy policy option implemented (with none, most-allocated, and least-allocated values)
  • Per-resource weights (numa-score-weights) implemented
  • kubelet_topology_manager_numa_score_selection_total metric implemented to help validate the feature is operative
  • Add proper e2e node tests

Alpha to Beta Graduation

  • Gather feedback from consumers of the new policy options
  • No major bugs reported in the previous cycle
  • Score-aware preferred-first merge optimization implemented
  • kubelet_topology_manager_numa_container_count metric implemented to track container distribution across NUMA nodes

Beta to GA Graduation

  • Allowing time for feedback (at least 2 releases as beta)
  • Risks have been addressed

Graduation Criteria of Options

This KEP follows the same option graduation model introduced in KEP-3545 and also used by CPUManagerPolicyOptions :

  • Alpha: Options are hidden by default and require both TopologyManagerPolicyOptions and TopologyManagerPolicyAlphaOptions feature gates to be enabled.
  • Beta: Options move to TopologyManagerPolicyBetaOptions and are available when TopologyManagerPolicyOptions is enabled (which is on by default since KEP-3545 graduated).
  • GA: Options become stable and are always available when TopologyManagerPolicyOptions is enabled.

Upgrade / Downgrade Strategy

Upgrade:

  • Feature gate default-off; users opt-in via kubelet config.
  • Score field added to TopologyHint with zero-value default. No behavior change for existing workloads.

Downgrade:

  • Rolling back kubelet removes policy options from config.
  • Pods already placed remain where they are.
  • New pods revert to Narrowest/Closest selection.

No changes to existing cluster configurations are required on upgrade.

Version Skew Strategy

Not applicable. This is a node-local kubelet feature with no cross-component dependencies. The feature does not affect the API server, scheduler, or any other control plane component.

Production Readiness Review Questionnaire

Feature Enablement and Rollback

How can this feature be enabled / disabled in a live cluster?
  • Feature gate (also fill in values in kep.yaml)
    • Feature gate names:
      • TopologyManagerPolicyOptions
      • TopologyManagerPolicyAlphaOptions
    • Components depending on the feature gate: kubelet
  • Change the kubelet configuration to set a TopologyManager policy (best-effort, restricted, or single-numa-node) and add numa-allocation-strategy (with value most-allocated or least-allocated) in TopologyManagerPolicyOptions.
    • Will enabling / disabling the feature require downtime of the control plane? No.
    • Will enabling / disabling the feature require downtime or reprovisioning of a node? Yes, a kubelet restart is required.
Does enabling the feature change any default behavior?

No. The policy options are opt-in. When numa-allocation-strategy is unset or "none", behavior is identical to the current Topology Manager behavior.

Can the feature be disabled once it has been enabled (i.e. can we roll back the enablement)?

Yes. Either:

  • Disable the TopologyManagerPolicyAlphaOptions feature gate, or
  • Set numa-allocation-strategy to "none", or
  • Remove the policy option from kubelet configuration.

In all cases, a kubelet restart is required. Existing pod placements are not affected; only new pod admissions will revert to the default Narrowest/Closest selection.

What happens if we reenable the feature if it was previously rolled back?

No impact on existing containers. The allocation-aware policy option will apply to new container admissions only.

Are there any tests for feature enablement/disablement?

There will be specific unit and e2e tests demonstrating that:

  • Default behavior is preserved when the feature gate is disabled.
  • Policy options are rejected when the feature gate is disabled.
  • Score field is ignored in hint comparison when the feature gate is disabled.

Rollout, Upgrade and Rollback Planning

How can a rollout or rollback fail? Can it impact already running workloads?

Kubelet may fail to start if invalid policy option values are provided (e.g., an unrecognized value for numa-allocation-strategy). Already running workloads are not affected. The feature only influences new pod admissions.

What specific metrics should inform a rollback?

A rise in kubelet_topology_manager_admission_errors_total (an existing kubelet metric) relative to kubelet_topology_manager_admission_requests_total, or unexpected pod placement patterns (e.g., pods not landing on expected NUMA nodes), should prompt investigation.

Alpha: The following metric will help operators confirm the feature is operative and identify unexpected placement patterns:

  • kubelet_topology_manager_numa_score_selection_total (counter): Tracks the total count of NUMA node selections where scoring influenced the placement decision (incremented when the selected hint differs from what the pre-existing narrowest/closest tiebreak would have chosen). Labels: numa_allocation_strategy (most-allocated, least-allocated). When this counter increases, it confirms the policy option is actively affecting placement decisions.

Operators can also manually verify placement by inspecting container CPU affinity via taskset -cp 1 and NUMA memory allocation via numactl -H inside containers.

Beta: An additional metric will be introduced, described under Monitoring Requirements :

  • kubelet_topology_manager_numa_container_count: tracks container distribution across NUMA nodes
Were upgrade and rollback tested? Was the upgrade->downgrade->upgrade path tested?

Will be tested as part of beta graduation criteria.

Is the rollout accompanied by any deprecations and/or removals of features, APIs, fields of API types, flags, etc.?

No.

Monitoring Requirements

How can an operator determine if the feature is in use by workloads?

Inspect the kubelet configuration of the nodes: check the feature gate status and the numa-allocation-strategy value in topologyManagerPolicyOptions.

A non-zero kubelet_topology_manager_numa_score_selection_total on a node confirms that the option is not merely configured but is actually changing placement decisions, and its numa_allocation_strategy label reports which strategy is in effect.

How can someone using this feature know that it is working for their instance?
  • Metrics
    • Metric name: kubelet_topology_manager_numa_score_selection_total
    • Components exposing the metric: kubelet

The counter increases whenever score-based selection picks a different NUMA node than the pre-existing narrowest/closest tiebreak would have picked, so a rising value is direct evidence that the feature is operative. A counter that stays at zero while the option is configured means every admission so far was already structurally determined (or no scored candidates tied), not necessarily that the feature is broken.

The metric is node-scoped and aggregate: it tells an operator that scoring is changing placement decisions on a node, but it does not attribute a decision to a particular pod, so a workload owner cannot use it alone to tell whether their pod was placed differently because of this feature. Per-workload attribution is a pre-existing gap in Topology Manager observability rather than one introduced here: none of the existing Topology Manager metrics or policy options report which hint was chosen for a given pod, and the admission decision is not surfaced in pod status. This KEP does not attempt to close that gap, but it also does not widen it, and any general solution (an admission event, or extending the PodResources API to report the selected affinity) would cover this option along with the rest.

  • Other (treat as last resort)
    • Details: To confirm placement for a specific workload, read the concrete assignments for its containers from the PodResources API List endpoint (KEP-2043 ), which reports the CPU IDs, memory regions, and devices assigned to each container along with their NUMA topology. This can be compared against the node’s allocation state at admission time to confirm the expected NUMA node was selected. Equivalently, launch a pod requiring resources from a specific NUMA node on a node with known asymmetric allocation and verify via taskset -cp 1 and numactl -H inside the container that resources are assigned from the expected NUMA node (most-allocated or least-allocated depending on the configured option).
What are the reasonable SLOs (Service Level Objectives) for the enhancement?

N/A. This is a node-local admission-time optimization. It does not introduce new API latency or availability concerns.

What are the SLIs (Service Level Indicators) an operator can use to determine the health of the service?
  • Metrics
    • Metric name: kubelet_topology_manager_admission_duration_ms
    • Components exposing the metric: kubelet

This existing metric captures the time spent in Topology Manager admission. Any regression introduced by score computation would be visible as an increase in this metric.

Are there any missing metrics that would be useful to have to improve observability of this feature?

Alpha: The following metric will be implemented to help validate that the feature is operative:

  • kubelet_topology_manager_numa_score_selection_total (counter): Total count of NUMA node selections where scoring influenced the placement decision (incremented when the selected hint differs from what the pre-existing narrowest/closest tiebreak would have chosen). Labels: numa_allocation_strategy (most-allocated, least-allocated). This helps operators verify that the policy option is taking effect and quantify its impact on placement decisions.

    Computing this does not require a second comparison pass: the merge loop already carries the incumbent best hint, which is exactly what Step 4 of Allocation-Aware Policy Options (“keep current”) selects under numa-allocation-strategy: none, so the two candidates are both available in the same pass. The counter is incremented once per completed merge, i.e. once per admitted container under container scope and once per admitted pod under pod scope.

Beta: The following metric will be added to improve observability:

  • kubelet_topology_manager_numa_container_count (gauge): Number of containers currently allocated on each NUMA node. Labels: numa_node. This metric helps operators verify expected pod placement patterns (e.g., containers spreading across NUMA nodes with least-allocated or consolidating with most-allocated) and identify NUMA imbalance issues.

    A container aligned to a multi-node affinity mask increments the gauge for every node in its mask, so the sum across nodes can exceed the number of running containers. The gauge is decremented when a container’s allocation is released, and is rebuilt from the checkpointed allocation state when the kubelet restarts.

Neither metric attributes a placement decision to an individual pod, so a workload owner still cannot tell from metrics alone whether their pod was placed differently because of this option. Per-pod attribution is a gap shared by all Topology Manager policy options, and a per-pod counter or gauge is the wrong shape for it: the cardinality is unbounded and the information is a one-shot admission fact, not a time series. At beta we will revisit this together with the wider Topology Manager observability discussion in SIG Node, where the plausible carriers are an admission-time event on the pod or the PodResources API reporting the affinity the Topology Manager selected. Both are broader than this KEP and are deliberately left out of scope for alpha.

Both metrics are registered in pkg/kubelet/metrics alongside the existing Topology Manager metrics, use the same kubelet subsystem, and are registered at ALPHA stability level.

Dependencies

Does this feature depend on any specific services running in the cluster?

No. The feature is entirely within the kubelet and uses allocation state already tracked by CPU manager, memory manager, and device manager.

Scalability

Will enabling / using this feature result in any new API calls?

No.

Will enabling / using this feature result in introducing new API types?

No.

Will enabling / using this feature result in any new calls to the cloud provider?

No.

Will enabling / using this feature result in increasing size or count of the existing API objects?

No.

Will enabling / using this feature result in increasing time taken by any operations covered by existing SLIs/SLOs?

No. The merge optimization (mergePreferred()) is expected to reduce hint merge time for common cases. Score computation adds O(n) overhead where n is the number of providers (typically 2-4), which is negligible.

Will enabling / using this feature result in non-negligible increase of resource usage (CPU, RAM, disk, IO, …) in any components?

No. The additional Score field in TopologyHint adds 8 bytes per hint, and the hint count is bounded and predictable. Hints are produced per requested resource by O(number of device plugins) + 2 providers (the CPU and memory managers), and each provider enumerates at most one hint per NUMA affinity mask it considers (bitmask.IterateBitMasks over N NUMA nodes). On real hardware (typically 2-8 NUMA nodes and 2-4 hint providers) this stays well under control, and the hints are transient: they are freed once admission completes. Score computation reuses allocation state already maintained by hint providers.

Can enabling / using this feature result in resource exhaustion of some node resources (PIDs, sockets, inodes, etc.)?

No.

Troubleshooting

How does this feature react if the API server and/or etcd is unavailable?

N/A. This is a kubelet-local feature that does not depend on the API server or etcd for its operation. Kubelet will continue to make admission decisions using local allocation state.

What are other known failure modes?
  • Invalid numa-allocation-strategy value (e.g., unrecognized string):

    • Detection: Kubelet fails to start with a clear error message listing valid values (none, most-allocated, least-allocated).
    • Diagnostics: Kubelet startup log will contain the validation error.
    • Testing: Unit tests cover value validation.
  • Invalid numa-score-weights values (out of range [0, 100], fractional, NaN):

    • Detection: Kubelet fails to start with a clear error message.
    • Mitigations: Fix the weight values in kubelet configuration. Weights must be integers in the range [0, 100], matching kube-scheduler’s NodeResourcesFit plugin.
    • Diagnostics: Kubelet startup log will contain the validation error.
    • Testing: Unit tests cover weight validation.
  • Containers of a pod land on different NUMA nodes under least-allocated:

    • Detection: Expected behavior under topologyManagerScope: "container" rather than a defect. See Interaction with Topology Manager Scope .
    • Mitigations: See Preserving Current Placement Behavior .
    • Diagnostics: Compare topologyManagerScope in the kubelet configuration against the pod’s observed CPU and memory affinity. The kubelet_topology_manager_numa_score_selection_total metric (alpha) confirms whether scoring is influencing placement decisions. In beta, the kubelet_topology_manager_numa_container_count metric will help identify when containers are spreading across NUMA nodes.
    • Testing: E2e tests cover both scopes.
What steps should be taken if SLOs are not being met to determine the problem?

N/A.

Implementation History

  • 2026-09-08: KEP created

Drawbacks

Adds complexity to the Topology Manager hint comparison and merge logic. However, the additional complexity is gated behind opt-in policy options and has no effect on the default code path. The change is well contained and clearly scoped, even at the projected GA stage.

Alternatives

Alternative 1: Scheduler-Based NUMA Balancing

Move NUMA allocation decisions to the scheduler instead of the kubelet.

Pros: Cluster-wide visibility, better global optimization. Cons: Requires scheduler-kubelet protocol changes; breaks existing Topology Manager design; significantly larger scope.

Alternative 2: Pod-Level Annotations

Let users specify packing/spreading preference per pod via annotations.

Pros: More flexible, users opt-in per workload. Cons: Increases configuration burden; hard to enforce cluster-wide policies; inconsistent with the existing node-level Topology Manager model. Annotations are also an API side channel: their values are not covered by API validation, so a malformed preference is only caught on the node at admission time, per node, rather than being rejected when the pod is created.

Alternative 3: Static NUMA Assignment

Pre-assign NUMA nodes to pod QoS classes.

Pros: Simple, predictable. Cons: Inflexible, wastes resources when QoS classes don’t match NUMA topology.

Alternative 4: Express This Through DRA and the CPU DRA Driver

Model CPUs as DRA devices with the CPU DRA driver and let the scheduler pick the NUMA domain when it allocates the claim, rather than adding a placement preference to the Topology Manager.

DRA already covers part of this. With KEP-6072 standardizing resource.kubernetes.io/numaNode and a matchAttribute constraint over it, a claim can require that its CPU, memory, GPU, and NIC devices come from the same NUMA node, which is the alignment guarantee the Topology Manager provides today. DRA is also the better long-term home for this: the scheduler sees allocation state cluster-wide in ResourceSlices, so it can avoid placing a pod on a node whose NUMA domains cannot satisfy it, instead of discovering that at admission time and rejecting the pod.

What is missing for this KEP’s use case:

  • DRA allocation satisfies constraints; it does not rank candidates. matchAttribute answers “may these devices be allocated together”, not “which of the several NUMA domains that all satisfy the constraint should be used”. There is no packing or spreading preference over equally valid device sets, which is exactly the decision this KEP addresses. Even enforcement: preferred on matchAttribute, which KEP-6072 lists as a non-goal, would soften the alignment requirement rather than order the candidates by utilization. A DRA equivalent needs device-level scoring in the scheduler, which does not exist today and is a considerably larger change than a Topology Manager policy option.
  • Not all the relevant resources are DRA devices. Memory and hugepages are managed by the memory manager, and CPUs are managed by the CPU manager for any pod not using the CPU DRA driver. Reconciling DRA-managed and node-allocatable resources is itself open work (KEP-5517 ). Until that is settled, a NUMA placement policy that spans CPU, memory, and devices has to live where all three are visible, which is the Topology Manager.
  • The workloads that need this run on the existing stack. The telco, NFV, and ML users driving kubernetes/kubernetes#125453 run the static CPU manager policy with the Topology Manager today, and will for several releases. Requiring them to migrate their resource management model to obtain a placement preference is disproportionate to the size of the change.

This KEP is therefore complementary rather than competing, and it does not foreclose the DRA path. Score is an internal field on TopologyHint and the options are opt-in kubelet configuration, so nothing here becomes API surface that a future DRA-based mechanism would have to carry. If CPU and memory allocation eventually move to DRA, the natural place for a utilization preference is the DRA allocator, and these options would be deprecated along with the rest of the Topology Manager policy options rather than separately.

Infrastructure Needed (Optional)

  • E2e test infrastructure with multi-NUMA nodes (at least 2 NUMA nodes per test worker).
  • Performance test cluster for benchmarking overhead.