Pidoku

The Frontend: From JSON to Token IDs

Basic 45 min Difficulty 2/5 Lesson 02 of 04

Prerequisites Processes and Wires

The idea in one minute#

Everything that happens before the engine core sees a request happens in the API server process, on the CPU. A JSON body is validated, a chat template turns a list of messages into one string, a tokenizer turns the string into integers, limits are checked, and finally an admission check decides whether the server has room at all. None of this uses the GPU, and all of it adds to the time to first token. On long prompts and with images it can be the largest part of that time, which is why it has its own process, its own thread pools and, increasingly, its own machines.

A picture#

flowchart LR
  J[":i-file-text: <b>JSON body</b><br/><small>messages, tools, sampling</small>"] --> V[":pydantic: <b>Validate</b><br/><small>request schema</small>"]
  V --> T[":i-code: <b>Chat template</b><br/><small>Jinja2 to one string</small>"]
  T --> K[":huggingface: <b>Tokenizer</b><br/><small>thread pool</small>"]
  IMG[":i-eye: <b>Images, audio</b><br/><small>fetch, decode, preprocess</small>"] --> MM[":i-cpu: <b>MM processor</b><br/><small>single worker + cache</small>"]
  K --> IP[":i-funnel: <b>InputProcessor</b><br/><small>length, params, LoRA</small>"]
  MM --> IP
  IP --> AD[":i-shield-check: <b>Admission</b><br/><small>queue limits to 503</small>"]
  AD --> E[":vllm: <b>Engine core</b><br/><small>EngineCoreRequest</small>"]
  class J,IMG neutral
  class V,T,K,MM compute
  class IP,AD queue
  class E io

How it really works#

The server#

The API server is a FastAPI application run by uvicorn on a single asyncio event loop. Route handlers live under vllm/entrypoints/, one sub-package per API family: openai/, anthropic/, cohere/, pooling/, speech_to_text/. All of them end by calling the same method, AsyncLLM.generate() (or .encode() for embedding models).

Because there is one event loop, anything slow must be moved off it or every other client waits. That rule explains the structure of this whole lesson.

warning

--api-key protects only paths under /v1, /v2 and /inference. Other routes on the same port — /invocations among them, which can run inference — are not authenticated. Operational routes such as /metrics, /health, /tokenize and, when enabled, the development endpoints are open too. Do not expose the port directly to an untrusted network; put a reverse proxy or gateway in front and allow only the paths you intend to serve.

Step 1: the chat template#

A chat model was trained on text with a precise layout: special tokens marking where each turn begins and ends and who is speaking. The chat template is a Jinja2 program, shipped in the model’s tokenizer configuration, that produces that layout from a list of messages.

messages:  [{"role": "system", "content": "Be brief."},
            {"role": "user",   "content": "Hi"}]

rendered:  <|im_start|>system\nBe brief.<|im_end|>\n
           <|im_start|>user\nHi<|im_end|>\n
           <|im_start|>assistant\n

The trailing <|im_start|>assistant\n is the generation prompt: it tells the model that the next thing to write is its own reply.

Things worth knowing about templates in vLLM:

  • No template, no chat. If the model ships none and you do not pass --chat-template, every /v1/chat/completions request fails. /v1/completions still works; it sends your string as written.
  • Templates take arguments. chat_template_kwargs in the request body is passed through. This is how per-request switches such as “enable thinking” reach models that support them.
  • Tools go through the template. The tools array is rendered into the prompt by the template; there is no separate channel. See Tool Calling and Reasoning.
  • The template is treated as untrusted code. It comes from a model repository, so it runs in Jinja2’s sandboxed environment with a time limit (VLLM_CHAT_TEMPLATE_RENDER_TIMEOUT); a template that hangs is abandoned and its worker thread replaced.

The class that owns templating and tokenising is the renderer (vllm/renderers/). HfRenderer handles anything with a Hugging Face tokenizer; a few model families (mistral.py, deepseek_v4.py and others) have their own.

Step 2: the tokenizer#

Tokenising is CPU work that can take tens of milliseconds for a long prompt. Doing it on the event loop would freeze every other connection, so the renderer hands it to a thread pool:

Python
# Thread pool executor for blocking tokenizer operations.
pool_workers = config.model_config.renderer_num_workers
self._executor = _SwappableExecutor(max_workers=pool_workers)

# Separate single-worker executor so tokenization never queues behind
# MM preprocessing; must stay single-worker (P0/P1 order).
self._mm_executor = ThreadPoolExecutor(max_workers=1)

renderer_num_workers defaults to 1. Hugging Face “fast” tokenizers are written in Rust and release the interpreter lock while they work, so one thread is enough to keep the loop responsive; raise it when many long prompts arrive together.

Multimodal preprocessing — resizing images, extracting video frames, computing audio features — gets a separate single-thread pool. Text requests therefore never wait behind an image. The preprocessed result is cached by content hash, so the same image sent twice is processed once; that cache is mirrored in the engine core so the pixels themselves do not cross the socket a second time (Multimodal).

Step 3: InputProcessor#

vllm/v1/engine/input_processor.py turns a rendered prompt plus parameters into an EngineCoreRequest. It is where most client mistakes are caught, and they are caught before any GPU time is spent:

CheckRuleWhat the client sees
Empty promptZero prompt tokens is an error for a generative model400
Prompt too longprompt_len > max_model_len400: “The decoder prompt (length N) is longer than the maximum model length of M”
No room to answerprompt_len == max_model_len400: prompt “plus the number of requested output tokens (at least 1) is longer than…”
max_tokens unsetIt is set to max_model_len - prompt_lenThe answer may run to the end of the context window
Token IDs out of rangeAn ID not in the vocabulary400
LoRA not enabledThe request names an adapter but the server was started without LoRA support400: “Got lora_request … but LoRA is not enabled!”
Sampling parametersTypes and ranges, and that any structured-output schema is acceptable to the backend400

One line here deserves attention:

Python
request.request_id = f"{request.external_req_id}-{random_uuid():.8}"

Whatever request ID the client supplied, the engine appends eight random characters. Two clients sending the same X-Request-Id therefore cannot collide inside the scheduler. The original ID is kept as external_req_id and is what appears in responses.

Step 4: n > 1 becomes n requests#

If the client asks for n: 4, the engine has no notion of “one request with four answers”. AsyncLLM.add_request creates a ParentRequest and four child requests with their own IDs and seeds. They are scheduled as four independent sequences. Their identical prompt is computed once and shared through prefix caching, and the parent reassembles the outputs. For admission purposes, a request with n = 4 counts as four.

Step 5: admission control#

By default vLLM’s waiting queue is unbounded: requests are accepted until memory runs out, and each waits as long as it must. Under overload that is the wrong behaviour — latency grows without limit and clients time out after the work has already been queued. Two optional limits turn excess load into an immediate, retryable error:

FlagCountsBehaviour at the limit
--max-num-queued-reqsUnfinished requests, waiting and running, across all engines this API server routes toNew requests get HTTP 503
--max-num-queued-tokensTotal prompt tokens of requests that have not yet produced their first tokenNew requests get HTTP 503

503 is deliberate: load balancers and client SDKs treat it as “try another instance”.

The token limit is a time-to-first-token guard, and the config’s own docstring gives the formula:

max_num_queued_tokens = target_TTFT × prefill_throughput

e.g. target 2 s, measured prefill speed 20,000 tokens/s  →  40,000

If 40,000 prompt tokens are already waiting to be read, a new request cannot possibly start answering within two seconds, so it is better refused than accepted.

The count is deliberately pessimistic. A request contributes its whole prompt length until its first output token arrives, even if most of the prompt has been read or was found in the prefix cache, because the engine core reports nothing to the API server during prefill. The source says so plainly: this “overestimates the real backlog, causing earlier rejection than strictly necessary — the safe direction for QoS.”

With several API server processes, the request count is kept in a small shared-memory array (SharedAdmissionStats) so that all of them see the same total.

Step 6: hand-off#

AsyncLLM._add_request then does three things in order:

Python
self.check_admission(request_id=request.request_id)
self.output_processor.add_request(request, prompt, parent_req, index, queue)
await self.engine_core.add_request_async(request)

The output side is registered before the request is sent. If it were the other way round, a very fast engine could return the first token for a request the API server did not yet know about.

Splitting the frontend off#

Since the engine core only ever sees token IDs, the frontend can run somewhere else:

  • vllm launch render <model> starts a server with no GPU that exposes /render and /derender (text to tokens and back).
  • vllm serve --enable-scale-out adds a tokens-in, tokens-out endpoint (/inference/v1/generate) to an ordinary server.
  • An experimental API server written in Rust (VLLM_USE_RUST_FRONTEND=1) replaces the Python one and speaks the same socket protocol to the engine core (The Rust Frontend and What Is Next).

A router that knows the token IDs before choosing a replica can send a request to the replica whose prefix cache already holds that prompt. That is the reason the split exists.

Code#

Size the two admission limits, then watch them work. The simulation feeds a server more prompt tokens per second than it can read and compares an unbounded queue with a bounded one.

Go
package main

import "fmt"

func main() {
	const (
		prefillSpeed = 20000.0 // prompt tokens the GPU reads per second
		targetTTFT   = 2.0     // seconds
		promptTokens = 4000    // every request, for simplicity
		arrivalRate  = 7.0     // requests per second: 28,000 tokens/s, 40% over capacity
		seconds      = 30
	)
	maxQueuedTokens := int(targetTTFT * prefillSpeed)
	fmt.Printf("max_num_queued_tokens = %.0f s x %.0f tok/s = %d\n\n",
		targetTTFT, prefillSpeed, maxQueuedTokens)

	run := func(limit int) (accepted, rejected int, worstWait float64) {
		backlog := 0.0 // prompt tokens waiting to be read
		const tick = 0.01
		carry := 0.0
		for t := 0.0; t < seconds; t += tick {
			backlog -= prefillSpeed * tick
			if backlog < 0 {
				backlog = 0
			}
			carry += arrivalRate * tick
			for carry >= 1 {
				carry--
				if limit > 0 && int(backlog) >= limit {
					rejected++
					continue
				}
				// This request waits for everything ahead of it, then its own prompt.
				wait := (backlog + promptTokens) / prefillSpeed
				if wait > worstWait {
					worstWait = wait
				}
				backlog += promptTokens
				accepted++
			}
		}
		return
	}

	for _, limit := range []int{0, maxQueuedTokens} {
		a, r, w := run(limit)
		name := "unbounded queue"
		if limit > 0 {
			name = fmt.Sprintf("limit %d tokens", limit)
		}
		fmt.Printf("%-20s accepted %3d  rejected(503) %3d  worst time-to-first-token %.1f s\n",
			name, a, r, w)
	}
}

Without the limit every request is accepted and the last ones wait many seconds for their first token. With it, about a quarter are refused instantly and everyone who is accepted is served close to the target. Refusing is the kinder behaviour: a 503 in one millisecond can be retried elsewhere; a twelve-second wait cannot be taken back.

Remember this#

  • The frontend is CPU work in the API server: validate, template, tokenise, check, admit.
  • A chat template is a Jinja2 program from the model repository; no template means no chat API.
  • Tokenising runs on a thread pool; image and audio preprocessing on a separate one.
  • InputProcessor rejects bad requests before any GPU time is spent, and randomises request IDs.
  • n > 1 is n independent requests that share a prompt through the prefix cache.
  • The queue is unbounded unless you set --max-num-queued-reqs or --max-num-queued-tokens; both answer 503.
  • --api-key does not cover every route.

Try it#

  1. Send a chat request with "max_tokens" omitted and a tiny prompt. Look at usage in the response and at the server’s max_model_len. What was max_tokens set to?
  2. Call /tokenize with a chat messages body, then /detokenize with the IDs it returns. The text you get back is the rendered chat template. Find the generation prompt at the end.
  3. In the program, set arrivalRate to 4.5 (below capacity). How many requests are rejected with the limit on? What does that tell you about the cost of enabling the limit when the server is healthy?

Check yourself#

  1. Why does tokenising run on a thread pool instead of directly in the request handler?
  2. A request has a prompt exactly as long as max_model_len. Is it accepted? Why?
  3. Why does the queued-token count use the full prompt length even when most of the prompt is cached?

Sources#

Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom