Topology-aware scheduling

Pin each replica group to a single topology domain, with native placement or Kueue.

AI inference workloads need constant pod-to-pod communication, so network bandwidth matters. The bandwidth between pods depends on how close their nodes sit in the data center. Topology-aware scheduling keeps each replica group within one topology domain to maximize it, which raises bandwidth for tensor and pipeline parallelism.

There are two ways to do this. Pick based on how your cluster admits workloads:

  • Native placement policies — LWS’s built-in leaderworkerset.sigs.k8s.io/exclusive-topology annotation. This is limited to exclusive 1:1 placement: each replica group (pod group) gets a topology domain to itself. No extra components. Use this when LWS schedules directly against the cluster.
  • Kueue — hand admission and placement to Kueue topology-aware scheduling (TAS). Unlike native exclusive-topology, Kueue can binpack more than one group into the same topology domain — the key thing Kueue enables here. Use this when Kueue already manages quota and admission for your cluster, so topology placement stays consistent with the rest of your gang scheduling.

Option 1: Native placement policies

This guide is the basic deployment plus topology-aware placement. The leaderworkerset.sigs.k8s.io/exclusive-topology annotation keeps each replica group within one topology domain and excludes other groups from it.

Set the annotation value to your cluster’s topology key (the example uses cloud.google.com/gke-nodepool). Nodes must be labeled with that key.

export HF_TOKEN=<your-hf-token>
curl https://raw.githubusercontent.com/kubernetes-sigs/lws/refs/heads/main/docs/examples/leaderworkerset/topology-aware-scheduling/vllm.yaml -s | envsubst | kubectl apply -f -
export HF_TOKEN=<your-hf-token>
curl https://raw.githubusercontent.com/kubernetes-sigs/lws/refs/heads/main/docs/examples/leaderworkerset/topology-aware-scheduling/sglang.yaml -s | envsubst | kubectl apply -f -

Option 2: Kueue topology-aware scheduling

Use this when Kueue already manages quota and admission for your cluster.

Kueue topology-aware scheduling (TAS) places the group through kueue.x-k8s.io/podset-required-topology. Unlike the native exclusive-topology feature, this is all-or-nothing: Kueue only admits the group if it fits entirely within one topology domain, and holds it otherwise. Also unlike native exclusive-topology — which reserves a domain for a single group — Kueue can binpack more than one group into the same topology domain. Prefer Kueue TAS when you want quota-gated, all-or-nothing placement with binpacking; prefer native placement when LWS schedules directly against the cluster.

Define topology levels

In a yaml file, define the levels of your topology and the resource type you schedule on.

kueuePopulator:
  config:
    topology:
      levels:
        - nodeLabel: "cloud.google.com/gce-topology-block"
        - nodeLabel: "cloud.google.com/gce-topology-subblock"
        - nodeLabel: "cloud.google.com/gce-topology-host"
        - nodeLabel: "kubernetes.io/hostname"
    resourceFlavor:
      nodeLabels:
        cloud.google.com/gke-gpu: "true"

Install the Kueue controller

helm install kueue oci://registry.k8s.io/kueue/charts/kueue --version=0.16.1 \
  --create-namespace --namespace=kueue-system

Now install kueue-populator, passing the topology definition:

helm install kueue-populator oci://registry.k8s.io/kueue/charts/kueue-populator \
  --version=0.16.1 --namespace=kueue-system --create-namespace --wait \
  -f <topology-yaml-file>

Deploy

export HF_TOKEN=<your-hf-token>
curl https://raw.githubusercontent.com/kubernetes-sigs/lws/refs/heads/main/docs/examples/leaderworkerset/topology-aware-scheduling/vllm-kueue.yaml -s | envsubst | kubectl apply -f -

See basic for how to reach the service once pods are running.

Feedback

Was this page helpful?