Pidoku

The Serving Layer

Intermediate 55 min Difficulty 3/5 Lesson 04 of 06

Prerequisites The Cluster, Model Access and Gateways

The idea in one minute#

The serving layer turns model weights into an endpoint that is fast, scalable and replaceable. It has three parts. The engine runs one model on one set of GPUs as efficiently as possible. The router decides which replica gets each request — and for language models that decision is worth as much as the engine, because a replica that already holds a request’s prefix in its cache answers several times sooner. The lifecycle controller adds and removes replicas and rolls out new versions. Design all three; most teams only think about the first.

A picture#

flowchart LR
  GW[":envoyproxy: <b>AI gateway</b><br/><small>auth, token limits</small>"] --> EPP
  subgraph ROUTE["Inference routing"]
    EPP[":llm-d: <b>Endpoint picker</b><br/><small>queue depth, KV use,<br/>prefix match, adapters</small>"]
  end
  subgraph POOL["InferencePool: one model"]
    direction TB
    R1[":vllm: <b>Replica 1</b><br/><small>prefix A cached</small>"]
    R2[":vllm: <b>Replica 2</b><br/><small>prefix B cached</small>"]
    R3[":vllm: <b>Replica 3</b><br/><small>warming</small>"]
  end
  EPP --> R1
  EPP --> R2
  EPP -.->|"not ready"| R3
  R1 -.->|"metrics every few seconds"| EPP
  R2 -.-> EPP
  AS[":keda: <b>Autoscaler</b><br/><small>on queue and KV pressure</small>"] --> POOL
  REG[(":huggingface: <b>Model registry</b><br/><small>signed, versioned</small>")] --> POOL
  R1 --> G1[(":nvidia: GPUs")]
  R2 --> G2[(":nvidia: GPUs")]
  class GW queue
  class EPP queue
  class R1,R2,R3 compute
  class AS neutral
  class REG,G1,G2 memory

How it really works#

The engine#

An inference engine does four things a naive loop does not, and together they are worth roughly a tenfold difference in cost:

MechanismWhat it doesDeep dive
Continuous batchingRequests join and leave a running batch at every decode step, so the GPU is never waiting for a batch to fillContinuous batching
KV cache with pagingKeeps attention state per sequence in fixed blocks, without fragmentationPagedAttention
Prefix cachingReuses the cached state of a shared prompt prefix across requestsPrefix caching
Structured decodingConstrains output to a schema during generationSampling

The engines in common use, as of October 2026:

EngineChoose it when
vLLMThe default: broadest model support, simplest deployment, largest community
SGLangWorkloads dominated by shared prefixes and structured output
TensorRT-LLMOne or two stable models at very high volume on NVIDIA hardware, where compiled engines repay the operational cost
llama.cpp, OllamaLaptops, edge and development — not shared serving

Text Generation Inference was archived in March 2026; do not start new work on it. Choosing a stack has the full decision tree. All of these expose the OpenAI-compatible API, which is what keeps the engine swappable.

The router#

Round-robin is wrong for language models for two reasons: requests differ a thousandfold in cost, and replicas differ in what they have cached. An inference-aware router scores replicas on live signals:

  • Queue depth and KV-cache utilisation — the real measures of load.
  • Prefix affinity — send a conversation back to the replica that holds its history.
  • Loaded adapters — which LoRA fine-tunes are already in memory.
  • Criticality — shed batch traffic before interactive.

On Kubernetes this is standardised as the Gateway API Inference Extension: an InferencePool groups the pods serving a model, and an endpoint picker chooses among them per request. Any conformant gateway can use it. llm-d (a CNCF sandbox project since March 2026) supplies a production endpoint picker with cache-aware scheduling; NVIDIA Dynamo (1.0 in 2026) provides its own router plus a KV-block manager that can move cache between nodes. They are layers above the engine, not replacements for it.

Disaggregation#

A request has two phases with opposite needs: prefill reads the prompt (compute-heavy, bursty) and decode generates tokens (memory-bandwidth-bound, steady). On one replica a long prefill stalls everyone’s decoding. Prefill/decode disaggregation runs them on separate pools and ships the KV cache between them. It pays off at scale with long prompts — exactly the agentic profile — and costs a fast network and real complexity. It is an optimisation to reach for after routing and caching are working, not before. Disaggregation has the details.

Autoscaling#

warning

Do not autoscale model servers on GPU utilisation. A GPU reads 100% utilised with one request or with sixty; the number says nothing about remaining capacity.

Scale on what users feel and what the engine reports: queue depth or queue wait time, KV-cache utilisation, and TTFT against its target. Scale out quickly and in slowly, because a wrong scale-in costs a cold start. And budget the cold start honestly:

new replica on a warm node      image cached, weights on NVMe      20 - 60 s
new replica on a cold node      pull image, pull weights, load     3 - 10 min
new node                        provision + driver + all above     5 - 15 min

So interactive pools keep warm spares, and scale-to-zero is for models that can afford a queue. Autoscaling and cold starts shows how to shorten each term.

Serving many models#

Few platforms serve one model. The usual mix:

  • A few large chat or coding models — dedicated pools, sized individually.
  • Fine-tuned variants — served as LoRA adapters on a shared base model, loaded on demand, so twenty customisations cost one model’s memory.
  • Small models — embeddings, rerankers, classifiers, guardrails — packed onto shared GPUs.
  • The long tail — rarely used models that scale to zero and accept a cold start.

Rolling out a model version#

A new model version, a new quantization or a new engine version is a behaviour change. Treat it like one:

  1. The artifact comes from the registry, signed and versioned; nothing is pulled by a floating tag.
  2. It passes the offline evaluation gate.
  3. It takes a canary share of traffic; quality signals and latency are compared with the baseline on live traffic.
  4. It is promoted gradually, with the old version kept warm until the new one is proven.
  5. Rollback is a routing change, not a redeploy.

See model rollouts.

What to hand to the platform layer#

The serving layer’s contract upward is small: an OpenAI-compatible endpoint per logical model, a health and readiness signal, and metrics — queue depth, KV utilisation, TTFT and TPOT histograms, tokens per second, per model and replica. Everything in the next two lessons is built on those.

Remember this#

  • Three parts: engine, router, lifecycle controller.
  • vLLM is the default engine; choose another for a stated reason.
  • Route on queue depth, KV-cache use and prefix affinity — never round-robin.
  • Autoscale on queue and cache pressure, never on GPU utilisation; keep warm spares.
  • A model version is rolled out like code: signed artifact, evaluation, canary, fast rollback.

Try it#

  1. Your agent workload has 40,000-token prompts of which 35,000 repeat between steps. Which router signal matters most, and what happens to latency with round-robin?
  2. A model takes 6 minutes to cold-start and traffic doubles in 2 minutes each morning. Write the scaling policy.
  3. List the models a platform at your company would need to serve. Sort them into the four groups above.

Check yourself#

  1. Why is the routing decision worth as much as the engine for language models?
  2. Why is GPU utilisation useless as an autoscaling signal?
  3. What makes rollback of a model version fast?

Sources#

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom