Pidoku

Deploying on Kubernetes

Expert 45 min Difficulty 3/5 Lesson 05 of 05

Prerequisites Your First Server, Metrics, Tuning and Benchmarking

The idea in one minute#

A vLLM pod is an unusual Kubernetes citizen. It takes minutes to start, it holds state that makes replicas unequal (each has a different prefix cache), it is saturated long before its CPU or memory graphs say so, and losing it discards every request in flight. A Deployment with default probes and a round-robin Service runs, and is wrong in each of those four ways. This lesson is the list of things to configure differently, the reason for each in terms of the internals you now know, and the projects that build the missing cluster-level layer.

A picture#

flowchart LR
  U[":i-users: Clients"] --> GW[":envoyproxy: <b>Gateway + inference router</b><br/><small>prefix- and load-aware</small>"]
  GW --> P1[":vllm: <b>vLLM pod</b><br/><small>prefix cache A</small>"]
  GW --> P2[":vllm: <b>vLLM pod</b><br/><small>prefix cache B</small>"]
  GW --> P3[":vllm: <b>vLLM pod</b><br/><small>starting: 3 min</small>"]
  P1 -.->|"/metrics, KV events"| GW
  P2 -.-> GW
  P1 --- PVC[(":i-hard-drive: <b>Shared volume</b><br/><small>model weights,<br/>compile cache</small>")]
  P2 --- PVC
  P3 --- PVC
  AS[":keda: <b>Autoscaler</b><br/><small>on queue depth, not GPU %</small>"] -.-> P3
  PR[":prometheus: Prometheus"] -.-> AS
  P1 -.-> PR
  class U neutral
  class GW queue
  class P1,P2 compute
  class P3 warn
  class PVC memory
  class AS,PR io

How it really works#

The pod#

A minimal correct container spec differs from a web service’s in five places.

1. GPU and CPU. Request the GPU, and give the pod enough CPU for its processes. From Processes and Wires: at least one physical core per process, so 2 + N for N GPUs. Set CPU requests equal to limits. A throttled engine core — whose busy loop is descheduled mid-step — shows up as lost throughput with an idle GPU, and CPU throttling is the cause people find last.

2. Shared memory. The engine core and workers communicate through a shared-memory ring buffer. A container’s default /dev/shm is 64 MiB, which is too small. Mount an in-memory volume:

YAML
volumes:
  - name: shm
    emptyDir: { medium: Memory, sizeLimit: "8Gi" }
containers:
  - volumeMounts:
      - { name: shm, mountPath: /dev/shm }

This is the Kubernetes equivalent of Docker’s --ipc=host. It is needed for tensor parallelism in particular.

3. Model storage. Downloading tens of gigabytes on every pod start is slow and can be rate-limited. Use a persistent volume mounted at the Hugging Face cache path, a pre-populated node-local disk, or a streaming loader that reads directly from object storage (vLLM documents several: Run:ai Model Streamer, Tensorizer, fastsafetensors). Provide the hub token from a Secret.

4. The compile cache. Mount a volume at /root/.cache/vllm, or bake the cache into the image. Otherwise every cold pod repeats torch.compile (CUDA Graphs and Compilation).

5. Probes. This is where default configurations fail.

Probes for a server that starts slowly#

vLLM does not listen on its port until the engine is fully initialised: weights loaded, memory profiled, cache allocated, model compiled, graphs captured. That takes from under a minute to many minutes. A liveness probe that starts checking after 30 seconds will kill a healthy pod that is still loading, and Kubernetes will restart it, forever.

The documentation describes the symptom: the container log ends with KeyboardInterrupt: terminated, and “the Kubernetes scheduler will kill the container” because “the startup or readiness probe failureThreshold is too low for the time needed to start up the server.”

Use three probes with distinct jobs:

YAML
startupProbe:                 # "has it finished starting?" — generous
  httpGet: { path: /health, port: 8000 }
  periodSeconds: 10
  failureThreshold: 90        # allow up to 15 minutes
livenessProbe:                # "is it dead?" — only after startup succeeded
  httpGet: { path: /health, port: 8000 }
  periodSeconds: 10
  failureThreshold: 3
readinessProbe:               # "should it receive traffic?"
  httpGet: { path: /health, port: 8000 }
  periodSeconds: 5

The startup probe holds the other two off until it passes. Measure your real startup time — the init engine … took line plus download and load — and set the threshold at two to three times that.

What /health means matters too. It returns 200 when the engine is up, and starts failing when the engine core has died (Processes and Wires: the ENGINE_CORE_DEAD path). It does not tell you the server is overloaded; a saturated vLLM is healthy. Do not use liveness to shed load.

Shutting down#

When a pod is deleted, Kubernetes sends SIGTERM and waits terminationGracePeriodSeconds, 30 by default. In-flight generations can run far longer than that.

Set the grace period to cover your longest expected response, and make sure traffic stops arriving first: remove the pod from the router, then let requests drain. vLLM has a drain path for this (wait_for_requests_to_drain, with a default limit of 300 seconds) and a shutdown state in which new requests are rejected while existing ones finish.

A rolling update therefore needs both: maxUnavailable: 0 so capacity never drops, and a grace period long enough that streams are not cut.

Why round-robin is the wrong load balancer#

Two properties of a vLLM replica break the assumption that replicas are interchangeable.

Each replica has its own prefix cache. A conversation’s second turn sent to a different replica than its first recomputes the whole history. With round-robin over R replicas, the chance of landing on the right one is 1/R: the more you scale out, the worse your cache hit rate becomes.

Requests are wildly unequal. One request is 50 tokens, the next 50,000. Counting requests or connections says nothing about load. A replica with three long generations may be full while another with thirty short ones is nearly idle.

A router that understands this uses what the pods publish:

SignalSourceUsed for
vllm:num_requests_waiting, num_requests_running/metricsAvoid replicas that are queueing
vllm:kv_cache_usage_perc/metricsAvoid replicas about to preempt
KV cache events (BlockStored, BlockRemoved)The KV event streamKnow which replica holds which prefix
The prompt’s block hashesComputed by the router from token IDsMatch a request to a replica’s cache

The last row is why vLLM has a GPU-less render server and a tokens-in endpoint (The Frontend), and why the default block hash uses a fixed seed: a router can compute, outside vLLM, the same hashes vLLM computes inside.

The cluster-level layer#

vLLM deliberately stops at one model per process. Several projects supply the rest; the documentation’s Kubernetes page lists more than a dozen integrations. The ones most tied to vLLM’s internals:

ProjectWhat it adds
vLLM production-stackA Helm chart from the vLLM organisation: a router with model-aware and prefix-aware routing, Grafana dashboards, KV offloading through LMCache
llm-dA CNCF Sandbox project (founded by Red Hat, Google Cloud, IBM Research, CoreWeave and NVIDIA) built around vLLM: prefix-aware routing from KV events, tiered cache offloading, prefill/decode disaggregation over the NIXL connector, wide expert parallelism, and autoscaling on inference signals
Gateway API Inference ExtensionThe Kubernetes standard for model-aware routing that such routers implement
AIBrixAlso from the vLLM organisation: LoRA management, autoscaling and distributed KV cache
KServe, KubeRay, LeaderWorkerSetGeneral model-serving and multi-node orchestration with vLLM as a runtime
NVIDIA DynamoA distributed serving framework that can use vLLM as its engine

llm-d’s own summary of the problem is a good statement of this whole lesson: “across many replicas, cache locality breaks under round-robin load balancing, long prompts inflate time-to-first-token, and accelerators sit underused.”

For multi-node single-replica deployments (tensor plus pipeline parallelism across machines), LeaderWorkerSet provides the “one leader pod plus N worker pods that start and stop together” primitive that a plain Deployment lacks.

Autoscaling#

GPU utilisation is the wrong signal. From Tuning and Benchmarking: a GPU serving a single request shows high utilisation, and one at its real limit shows the same. Scale on what saturation actually looks like:

SignalScale up when
vllm:num_requests_waitingSustained above a small number per replica
p95 of vllm:request_queue_time_secondsAbove a fraction of your TTFT target
vllm:kv_cache_usage_percSustained above about 0.9
Rate of 503 responsesAdmission control is rejecting

The other half of autoscaling is lead time. A new replica is useful only after it has started, and that takes minutes. Scaling from the current queue is therefore always late; headroom must cover the traffic that arrives while a pod boots. The program below computes it.

Ways to shorten the boot:

TechniqueWhat it removes
Weights on a fast local or shared volume, or a streaming loaderDownload
Persisted compile cachetorch.compile
-O1, or fewer CUDA graph sizesCapture time, at some run-time cost
Sleep mode (--enable-sleep-mode)Everything, for a pod kept alive: level 1 parks weights in host RAM and wakes in seconds (The Engine Core Loop)
vllm preloadWeight loading across restarts: a daemon per GPU keeps the sharded, quantised weights resident and hands them to a restarting engine without copying
Initialized snapshots (experimental)Initialisation itself: a checkpoint of the whole initialised engine is restored on the same machine

A deployment checklist#

  • CPU: at least 2 + N physical cores, requests equal to limits.
  • /dev/shm as an in-memory volume.
  • Weights and compile cache on persistent or pre-populated storage.
  • startupProbe sized to two or three times measured startup.
  • Grace period at least your longest response; maxUnavailable: 0.
  • A reverse proxy or gateway in front; --api-key does not protect every route (The Frontend). Never expose development endpoints.
  • --allowed-media-domains if the model accepts URLs.
  • Admission limits set from a benchmark.
  • Metrics scraped; alerts on queue time, preemptions and the engine-dead health failure.
  • Routing that is at least load-aware, and prefix-aware if conversations are long.
  • Autoscaling on queue signals, with headroom for boot time.

Code#

Probe thresholds and autoscaling headroom from measured numbers.

Go
package main

import (
	"fmt"
	"math"
)

func main() {
	// --- Measured for one pod (seconds) ---
	phases := []struct {
		name       string
		cold, warm float64 // warm = weights on local disk, compile cache present
	}{
		{"schedule pod + pull image", 45, 15},
		{"download weights", 240, 0},
		{"load weights to GPU", 40, 40},
		{"profile + create KV cache", 8, 8},
		{"torch.compile", 70, 6},
		{"capture CUDA graphs", 25, 25},
	}
	var cold, warm float64
	fmt.Printf("%-28s %8s %8s\n", "phase", "cold", "warm")
	for _, p := range phases {
		fmt.Printf("%-28s %7.0fs %7.0fs\n", p.name, p.cold, p.warm)
		cold += p.cold
		warm += p.warm
	}
	fmt.Printf("%-28s %7.0fs %7.0fs\n\n", "total", cold, warm)

	const period = 10.0
	for _, c := range []struct {
		name string
		t    float64
	}{{"cold", cold}, {"warm", warm}} {
		threshold := math.Ceil(2.5 * c.t / period)
		fmt.Printf("startupProbe (%s start): periodSeconds %.0f, failureThreshold %.0f  (allows %.0f s)\n",
			c.name, period, threshold, threshold*period)
	}

	// --- Autoscaling headroom ---
	const (
		perReplica = 12.0 // requests/s one replica sustains within the SLO (from a benchmark)
		current    = 60.0 // requests/s now
		growth     = 0.25 // traffic grows this fraction per minute during a ramp
	)
	fmt.Printf("\ntraffic %.0f req/s, growing %.0f%%/min, one replica handles %.0f req/s\n",
		current, growth*100, perReplica)
	fmt.Printf("%-12s %14s %18s %16s\n", "boot time", "traffic by then", "replicas needed", "spare to hold now")
	for _, boot := range []float64{warm, cold} {
		future := current * math.Pow(1+growth, boot/60)
		need := math.Ceil(future / perReplica)
		now := math.Ceil(current / perReplica)
		fmt.Printf("%9.0f s %12.0f r/s %18.0f %16.0f\n", boot, future, need, need-now)
	}
}

A cold start of seven minutes needs a startup probe that tolerates about eighteen. And during a ramp, the replicas you must already be running are determined by boot time: the cold-start row needs several spare replicas to stay within the SLO, the warm-start row far fewer. Every second removed from startup is capacity you no longer have to keep idle.

Remember this#

  • Give the pod 2 + N physical cores, an in-memory /dev/shm, and persistent weights and compile cache.
  • Use a startupProbe sized to real startup time; /health reports engine death, not overload.
  • Set the termination grace period to your longest response and drain before stopping.
  • Replicas are not interchangeable: each has its own prefix cache, and requests vary enormously in cost.
  • Route on queue depth, cache usage and prefix location, not round-robin.
  • Autoscale on waiting requests and queue time; keep headroom proportional to boot time.
  • Sleep mode, preload and a persisted compile cache each remove a part of cold start.
  • vLLM serves one model per process; production-stack, llm-d and others provide the fleet layer.

Try it#

  1. Replace the phase timings with numbers from your own cluster. What is your cold-to-warm ratio, and which phase dominates?
  2. With traffic growing 50% per minute, how many spare replicas does a cold start require? At what growth rate does scale-from-zero become impossible within your SLO?
  3. Deploy two replicas behind a round-robin Service and send a 20-turn conversation. Read cached_tokens on each turn. Repeat with session affinity on a conversation ID.

Check yourself#

  1. Why does a liveness probe with a 30-second initial delay put a vLLM pod into a restart loop?
  2. Why does adding replicas behind a round-robin balancer reduce the prefix-cache hit rate?
  3. Name two signals better than GPU utilisation for autoscaling, and say what each indicates.

Sources#

Checked on 5 October 2026.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom