Basic
Minimal multi-node inference on LeaderWorkerSet with vLLM and SGLang.
These guides deploy a distributed, multi-node inference service with
LeaderWorkerSet, spreading tensor and pipeline parallelism across the leader and
worker pods. Each guide isolates one feature and ships both a vllm.yaml and a
sglang.yaml.
exclusive-topology.Minimal multi-node inference on LeaderWorkerSet with vLLM and SGLang.
Scale LeaderWorkerSet replica groups with a HorizontalPodAutoscaler.
Pin each replica group to a single topology domain, with native placement or Kueue.
Was this page helpful?
Glad to hear it! Please tell us how we can improve.
Sorry to hear that. Please tell us how we can improve.