The idea in one minute#
A vLLM pod is an unusual Kubernetes citizen. It takes minutes to start, it holds state that makes replicas unequal (each has a different prefix cache), it is saturated long before its CPU or memory graphs say so, and losing it discards every request in flight. A Deployment with default probes and a round-robin Service runs, and is wrong in each of those four ways. This lesson is the list of things to configure differently, the reason for each in terms of the internals you now know, and the projects that build the missing cluster-level layer.
A picture#
flowchart LR
U[":i-users: Clients"] --> GW[":envoyproxy: <b>Gateway + inference router</b><br/><small>prefix- and load-aware</small>"]
GW --> P1[":vllm: <b>vLLM pod</b><br/><small>prefix cache A</small>"]
GW --> P2[":vllm: <b>vLLM pod</b><br/><small>prefix cache B</small>"]
GW --> P3[":vllm: <b>vLLM pod</b><br/><small>starting: 3 min</small>"]
P1 -.->|"/metrics, KV events"| GW
P2 -.-> GW
P1 --- PVC[(":i-hard-drive: <b>Shared volume</b><br/><small>model weights,<br/>compile cache</small>")]
P2 --- PVC
P3 --- PVC
AS[":keda: <b>Autoscaler</b><br/><small>on queue depth, not GPU %</small>"] -.-> P3
PR[":prometheus: Prometheus"] -.-> AS
P1 -.-> PR
class U neutral
class GW queue
class P1,P2 compute
class P3 warn
class PVC memory
class AS,PR ioHow it really works#
The pod#
A minimal correct container spec differs from a web service’s in five places.
1. GPU and CPU. Request the GPU, and give the pod enough CPU for its processes. From
Processes and Wires: at least one physical
core per process, so 2 + N for N GPUs. Set CPU requests equal to limits. A throttled engine
core — whose busy loop is descheduled mid-step — shows up as lost throughput with an idle GPU,
and CPU throttling is the cause people find last.
2. Shared memory. The engine core and workers communicate through a shared-memory ring
buffer. A container’s default /dev/shm is 64 MiB, which is too small. Mount an in-memory
volume:
volumes:
- name: shm
emptyDir: { medium: Memory, sizeLimit: "8Gi" }
containers:
- volumeMounts:
- { name: shm, mountPath: /dev/shm }This is the Kubernetes equivalent of Docker’s --ipc=host. It is needed for tensor
parallelism in particular.
3. Model storage. Downloading tens of gigabytes on every pod start is slow and can be rate-limited. Use a persistent volume mounted at the Hugging Face cache path, a pre-populated node-local disk, or a streaming loader that reads directly from object storage (vLLM documents several: Run:ai Model Streamer, Tensorizer, fastsafetensors). Provide the hub token from a Secret.
4. The compile cache. Mount a volume at /root/.cache/vllm, or bake the cache into the
image. Otherwise every cold pod repeats torch.compile
(CUDA Graphs and Compilation).
5. Probes. This is where default configurations fail.
Probes for a server that starts slowly#
vLLM does not listen on its port until the engine is fully initialised: weights loaded, memory profiled, cache allocated, model compiled, graphs captured. That takes from under a minute to many minutes. A liveness probe that starts checking after 30 seconds will kill a healthy pod that is still loading, and Kubernetes will restart it, forever.
The documentation describes the symptom: the container log ends with
KeyboardInterrupt: terminated, and “the Kubernetes scheduler will kill the container” because
“the startup or readiness probe failureThreshold is too low for the time needed to start up the
server.”
Use three probes with distinct jobs:
startupProbe: # "has it finished starting?" — generous
httpGet: { path: /health, port: 8000 }
periodSeconds: 10
failureThreshold: 90 # allow up to 15 minutes
livenessProbe: # "is it dead?" — only after startup succeeded
httpGet: { path: /health, port: 8000 }
periodSeconds: 10
failureThreshold: 3
readinessProbe: # "should it receive traffic?"
httpGet: { path: /health, port: 8000 }
periodSeconds: 5The startup probe holds the other two off until it passes. Measure your real startup time —
the init engine … took line plus download and load — and set the threshold at two to three
times that.
What /health means matters too. It returns 200 when the engine is up, and starts failing when
the engine core has died (Processes and Wires:
the ENGINE_CORE_DEAD path). It does not tell you the server is overloaded; a saturated
vLLM is healthy. Do not use liveness to shed load.
Shutting down#
When a pod is deleted, Kubernetes sends SIGTERM and waits terminationGracePeriodSeconds,
30 by default. In-flight generations can run far longer than that.
Set the grace period to cover your longest expected response, and make sure traffic stops
arriving first: remove the pod from the router, then let requests drain. vLLM has a drain path
for this (wait_for_requests_to_drain, with a default limit of 300 seconds) and a shutdown
state in which new requests are rejected while existing ones finish.
A rolling update therefore needs both: maxUnavailable: 0 so capacity never drops, and a grace
period long enough that streams are not cut.
Why round-robin is the wrong load balancer#
Two properties of a vLLM replica break the assumption that replicas are interchangeable.
Each replica has its own prefix cache. A conversation’s second turn sent to a different
replica than its first recomputes the whole history. With round-robin over R replicas, the
chance of landing on the right one is 1/R: the more you scale out, the worse your cache hit
rate becomes.
Requests are wildly unequal. One request is 50 tokens, the next 50,000. Counting requests or connections says nothing about load. A replica with three long generations may be full while another with thirty short ones is nearly idle.
A router that understands this uses what the pods publish:
| Signal | Source | Used for |
|---|---|---|
vllm:num_requests_waiting, num_requests_running | /metrics | Avoid replicas that are queueing |
vllm:kv_cache_usage_perc | /metrics | Avoid replicas about to preempt |
KV cache events (BlockStored, BlockRemoved) | The KV event stream | Know which replica holds which prefix |
| The prompt’s block hashes | Computed by the router from token IDs | Match a request to a replica’s cache |
The last row is why vLLM has a GPU-less render server and a tokens-in endpoint (The Frontend), and why the default block hash uses a fixed seed: a router can compute, outside vLLM, the same hashes vLLM computes inside.
The cluster-level layer#
vLLM deliberately stops at one model per process. Several projects supply the rest; the documentation’s Kubernetes page lists more than a dozen integrations. The ones most tied to vLLM’s internals:
| Project | What it adds |
|---|---|
| vLLM production-stack | A Helm chart from the vLLM organisation: a router with model-aware and prefix-aware routing, Grafana dashboards, KV offloading through LMCache |
| llm-d | A CNCF Sandbox project (founded by Red Hat, Google Cloud, IBM Research, CoreWeave and NVIDIA) built around vLLM: prefix-aware routing from KV events, tiered cache offloading, prefill/decode disaggregation over the NIXL connector, wide expert parallelism, and autoscaling on inference signals |
| Gateway API Inference Extension | The Kubernetes standard for model-aware routing that such routers implement |
| AIBrix | Also from the vLLM organisation: LoRA management, autoscaling and distributed KV cache |
| KServe, KubeRay, LeaderWorkerSet | General model-serving and multi-node orchestration with vLLM as a runtime |
| NVIDIA Dynamo | A distributed serving framework that can use vLLM as its engine |
llm-d’s own summary of the problem is a good statement of this whole lesson: “across many replicas, cache locality breaks under round-robin load balancing, long prompts inflate time-to-first-token, and accelerators sit underused.”
For multi-node single-replica deployments (tensor plus pipeline parallelism across machines), LeaderWorkerSet provides the “one leader pod plus N worker pods that start and stop together” primitive that a plain Deployment lacks.
Autoscaling#
GPU utilisation is the wrong signal. From Tuning and Benchmarking: a GPU serving a single request shows high utilisation, and one at its real limit shows the same. Scale on what saturation actually looks like:
| Signal | Scale up when |
|---|---|
vllm:num_requests_waiting | Sustained above a small number per replica |
p95 of vllm:request_queue_time_seconds | Above a fraction of your TTFT target |
vllm:kv_cache_usage_perc | Sustained above about 0.9 |
| Rate of 503 responses | Admission control is rejecting |
The other half of autoscaling is lead time. A new replica is useful only after it has started, and that takes minutes. Scaling from the current queue is therefore always late; headroom must cover the traffic that arrives while a pod boots. The program below computes it.
Ways to shorten the boot:
| Technique | What it removes |
|---|---|
| Weights on a fast local or shared volume, or a streaming loader | Download |
| Persisted compile cache | torch.compile |
-O1, or fewer CUDA graph sizes | Capture time, at some run-time cost |
Sleep mode (--enable-sleep-mode) | Everything, for a pod kept alive: level 1 parks weights in host RAM and wakes in seconds (The Engine Core Loop) |
vllm preload | Weight loading across restarts: a daemon per GPU keeps the sharded, quantised weights resident and hands them to a restarting engine without copying |
| Initialized snapshots (experimental) | Initialisation itself: a checkpoint of the whole initialised engine is restored on the same machine |
A deployment checklist#
- CPU: at least
2 + Nphysical cores, requests equal to limits. /dev/shmas an in-memory volume.- Weights and compile cache on persistent or pre-populated storage.
startupProbesized to two or three times measured startup.- Grace period at least your longest response;
maxUnavailable: 0. - A reverse proxy or gateway in front;
--api-keydoes not protect every route (The Frontend). Never expose development endpoints. --allowed-media-domainsif the model accepts URLs.- Admission limits set from a benchmark.
- Metrics scraped; alerts on queue time, preemptions and the engine-dead health failure.
- Routing that is at least load-aware, and prefix-aware if conversations are long.
- Autoscaling on queue signals, with headroom for boot time.
Code#
Probe thresholds and autoscaling headroom from measured numbers.
package main
import (
"fmt"
"math"
)
func main() {
// --- Measured for one pod (seconds) ---
phases := []struct {
name string
cold, warm float64 // warm = weights on local disk, compile cache present
}{
{"schedule pod + pull image", 45, 15},
{"download weights", 240, 0},
{"load weights to GPU", 40, 40},
{"profile + create KV cache", 8, 8},
{"torch.compile", 70, 6},
{"capture CUDA graphs", 25, 25},
}
var cold, warm float64
fmt.Printf("%-28s %8s %8s\n", "phase", "cold", "warm")
for _, p := range phases {
fmt.Printf("%-28s %7.0fs %7.0fs\n", p.name, p.cold, p.warm)
cold += p.cold
warm += p.warm
}
fmt.Printf("%-28s %7.0fs %7.0fs\n\n", "total", cold, warm)
const period = 10.0
for _, c := range []struct {
name string
t float64
}{{"cold", cold}, {"warm", warm}} {
threshold := math.Ceil(2.5 * c.t / period)
fmt.Printf("startupProbe (%s start): periodSeconds %.0f, failureThreshold %.0f (allows %.0f s)\n",
c.name, period, threshold, threshold*period)
}
// --- Autoscaling headroom ---
const (
perReplica = 12.0 // requests/s one replica sustains within the SLO (from a benchmark)
current = 60.0 // requests/s now
growth = 0.25 // traffic grows this fraction per minute during a ramp
)
fmt.Printf("\ntraffic %.0f req/s, growing %.0f%%/min, one replica handles %.0f req/s\n",
current, growth*100, perReplica)
fmt.Printf("%-12s %14s %18s %16s\n", "boot time", "traffic by then", "replicas needed", "spare to hold now")
for _, boot := range []float64{warm, cold} {
future := current * math.Pow(1+growth, boot/60)
need := math.Ceil(future / perReplica)
now := math.Ceil(current / perReplica)
fmt.Printf("%9.0f s %12.0f r/s %18.0f %16.0f\n", boot, future, need, need-now)
}
}A cold start of seven minutes needs a startup probe that tolerates about eighteen. And during a ramp, the replicas you must already be running are determined by boot time: the cold-start row needs several spare replicas to stay within the SLO, the warm-start row far fewer. Every second removed from startup is capacity you no longer have to keep idle.
Remember this#
- Give the pod
2 + Nphysical cores, an in-memory/dev/shm, and persistent weights and compile cache. - Use a
startupProbesized to real startup time;/healthreports engine death, not overload. - Set the termination grace period to your longest response and drain before stopping.
- Replicas are not interchangeable: each has its own prefix cache, and requests vary enormously in cost.
- Route on queue depth, cache usage and prefix location, not round-robin.
- Autoscale on waiting requests and queue time; keep headroom proportional to boot time.
- Sleep mode, preload and a persisted compile cache each remove a part of cold start.
- vLLM serves one model per process; production-stack, llm-d and others provide the fleet layer.
Try it#
- Replace the phase timings with numbers from your own cluster. What is your cold-to-warm ratio, and which phase dominates?
- With traffic growing 50% per minute, how many spare replicas does a cold start require? At what growth rate does scale-from-zero become impossible within your SLO?
- Deploy two replicas behind a round-robin Service and send a 20-turn conversation. Read
cached_tokenson each turn. Repeat with session affinity on a conversation ID.
Check yourself#
- Why does a liveness probe with a 30-second initial delay put a vLLM pod into a restart loop?
- Why does adding replicas behind a round-robin balancer reduce the prefix-cache hit rate?
- Name two signals better than GPU utilisation for autoscaling, and say what each indicates.
Sources#
Checked on 5 October 2026.
- Using Kubernetes — manifests, shared memory, the probe troubleshooting note
- vLLM production stack
- llm-d and llm-d.ai
- Sleep mode
- Preload
- Initialized engine snapshots
- Security guide