The idea in one minute#
The serving layer turns model weights into an endpoint that is fast, scalable and replaceable. It has three parts. The engine runs one model on one set of GPUs as efficiently as possible. The router decides which replica gets each request — and for language models that decision is worth as much as the engine, because a replica that already holds a request’s prefix in its cache answers several times sooner. The lifecycle controller adds and removes replicas and rolls out new versions. Design all three; most teams only think about the first.
A picture#
flowchart LR
GW[":envoyproxy: <b>AI gateway</b><br/><small>auth, token limits</small>"] --> EPP
subgraph ROUTE["Inference routing"]
EPP[":llm-d: <b>Endpoint picker</b><br/><small>queue depth, KV use,<br/>prefix match, adapters</small>"]
end
subgraph POOL["InferencePool: one model"]
direction TB
R1[":vllm: <b>Replica 1</b><br/><small>prefix A cached</small>"]
R2[":vllm: <b>Replica 2</b><br/><small>prefix B cached</small>"]
R3[":vllm: <b>Replica 3</b><br/><small>warming</small>"]
end
EPP --> R1
EPP --> R2
EPP -.->|"not ready"| R3
R1 -.->|"metrics every few seconds"| EPP
R2 -.-> EPP
AS[":keda: <b>Autoscaler</b><br/><small>on queue and KV pressure</small>"] --> POOL
REG[(":huggingface: <b>Model registry</b><br/><small>signed, versioned</small>")] --> POOL
R1 --> G1[(":nvidia: GPUs")]
R2 --> G2[(":nvidia: GPUs")]
class GW queue
class EPP queue
class R1,R2,R3 compute
class AS neutral
class REG,G1,G2 memoryHow it really works#
The engine#
An inference engine does four things a naive loop does not, and together they are worth roughly a tenfold difference in cost:
| Mechanism | What it does | Deep dive |
|---|---|---|
| Continuous batching | Requests join and leave a running batch at every decode step, so the GPU is never waiting for a batch to fill | Continuous batching |
| KV cache with paging | Keeps attention state per sequence in fixed blocks, without fragmentation | PagedAttention |
| Prefix caching | Reuses the cached state of a shared prompt prefix across requests | Prefix caching |
| Structured decoding | Constrains output to a schema during generation | Sampling |
The engines in common use, as of October 2026:
| Engine | Choose it when |
|---|---|
| vLLM | The default: broadest model support, simplest deployment, largest community |
| SGLang | Workloads dominated by shared prefixes and structured output |
| TensorRT-LLM | One or two stable models at very high volume on NVIDIA hardware, where compiled engines repay the operational cost |
| llama.cpp, Ollama | Laptops, edge and development — not shared serving |
Text Generation Inference was archived in March 2026; do not start new work on it. Choosing a stack has the full decision tree. All of these expose the OpenAI-compatible API, which is what keeps the engine swappable.
The router#
Round-robin is wrong for language models for two reasons: requests differ a thousandfold in cost, and replicas differ in what they have cached. An inference-aware router scores replicas on live signals:
- Queue depth and KV-cache utilisation — the real measures of load.
- Prefix affinity — send a conversation back to the replica that holds its history.
- Loaded adapters — which LoRA fine-tunes are already in memory.
- Criticality — shed batch traffic before interactive.
On Kubernetes this is standardised as the Gateway API Inference Extension: an
InferencePool groups the pods serving a model, and an endpoint picker chooses among them
per request. Any conformant gateway can use it. llm-d (a CNCF sandbox project since March
2026) supplies a production endpoint picker with cache-aware scheduling; NVIDIA Dynamo
(1.0 in 2026) provides its own router plus a KV-block manager that can move cache between
nodes. They are layers above the engine, not replacements for it.
Disaggregation#
A request has two phases with opposite needs: prefill reads the prompt (compute-heavy, bursty) and decode generates tokens (memory-bandwidth-bound, steady). On one replica a long prefill stalls everyone’s decoding. Prefill/decode disaggregation runs them on separate pools and ships the KV cache between them. It pays off at scale with long prompts — exactly the agentic profile — and costs a fast network and real complexity. It is an optimisation to reach for after routing and caching are working, not before. Disaggregation has the details.
Autoscaling#
warning
Do not autoscale model servers on GPU utilisation. A GPU reads 100% utilised with one request or with sixty; the number says nothing about remaining capacity.
Scale on what users feel and what the engine reports: queue depth or queue wait time, KV-cache utilisation, and TTFT against its target. Scale out quickly and in slowly, because a wrong scale-in costs a cold start. And budget the cold start honestly:
new replica on a warm node image cached, weights on NVMe 20 - 60 s
new replica on a cold node pull image, pull weights, load 3 - 10 min
new node provision + driver + all above 5 - 15 minSo interactive pools keep warm spares, and scale-to-zero is for models that can afford a queue. Autoscaling and cold starts shows how to shorten each term.
Serving many models#
Few platforms serve one model. The usual mix:
- A few large chat or coding models — dedicated pools, sized individually.
- Fine-tuned variants — served as LoRA adapters on a shared base model, loaded on demand, so twenty customisations cost one model’s memory.
- Small models — embeddings, rerankers, classifiers, guardrails — packed onto shared GPUs.
- The long tail — rarely used models that scale to zero and accept a cold start.
Rolling out a model version#
A new model version, a new quantization or a new engine version is a behaviour change. Treat it like one:
- The artifact comes from the registry, signed and versioned; nothing is pulled by a floating tag.
- It passes the offline evaluation gate.
- It takes a canary share of traffic; quality signals and latency are compared with the baseline on live traffic.
- It is promoted gradually, with the old version kept warm until the new one is proven.
- Rollback is a routing change, not a redeploy.
See model rollouts.
What to hand to the platform layer#
The serving layer’s contract upward is small: an OpenAI-compatible endpoint per logical model, a health and readiness signal, and metrics — queue depth, KV utilisation, TTFT and TPOT histograms, tokens per second, per model and replica. Everything in the next two lessons is built on those.
Remember this#
- Three parts: engine, router, lifecycle controller.
- vLLM is the default engine; choose another for a stated reason.
- Route on queue depth, KV-cache use and prefix affinity — never round-robin.
- Autoscale on queue and cache pressure, never on GPU utilisation; keep warm spares.
- A model version is rolled out like code: signed artifact, evaluation, canary, fast rollback.
Try it#
- Your agent workload has 40,000-token prompts of which 35,000 repeat between steps. Which router signal matters most, and what happens to latency with round-robin?
- A model takes 6 minutes to cold-start and traffic doubles in 2 minutes each morning. Write the scaling policy.
- List the models a platform at your company would need to serve. Sort them into the four groups above.
Check yourself#
- Why is the routing decision worth as much as the engine for language models?
- Why is GPU utilisation useless as an autoscaling signal?
- What makes rollback of a model version fast?