Pidoku

The Map: One Request, End to End

Foundations 45 min Difficulty 2/5 Lesson 03 of 03

Prerequisites Your First Server

The idea in one minute#

A request to vLLM crosses three processes and about a dozen named objects. Text becomes token IDs in the API server. Token IDs cross a socket to the engine core, where a scheduler decides — once per step — how many tokens of each request the GPU will process, and a KV cache manager hands out the memory blocks for them. The schedule crosses a second boundary to the worker, whose model runner builds tensors and calls the model. One new token per request comes back the same way, is turned into text, and is streamed to the client. Then the loop runs again. Everything else in this course is a closer look at one stop on this route.

A picture#

flowchart TB
  subgraph P0["Process 1 — API server (asyncio)"]
    direction LR
    H[":i-globe: <b>HTTP handler</b><br/><small>FastAPI route</small>"] --> R[":i-file-text: <b>Renderer</b><br/><small>chat template, tokenise</small>"]
    R --> IP[":i-funnel: <b>InputProcessor</b><br/><small>validate, build request</small>"]
    IP --> AL[":vllm: <b>AsyncLLM</b><br/><small>admission, per-request queue</small>"]
    OP[":i-message-square: <b>OutputProcessor</b><br/><small>detokenise, stop strings</small>"] --> H
  end
  subgraph P1["Process 2 — Engine core (busy loop)"]
    direction LR
    SCH[":i-list-checks: <b>Scheduler</b><br/><small>token budget, queues</small>"] --- KV[(":i-layers: <b>KVCacheManager</b><br/><small>block pool, prefix cache</small>")]
    SCH --> EX[":i-split: <b>Executor</b><br/><small>fan out to workers</small>"]
  end
  subgraph P2["Process 3..N — one worker per GPU"]
    direction LR
    MR[":pytorch: <b>Model runner</b><br/><small>build batch, sample</small>"] --> M[":nvidia: <b>Model + attention kernel</b>"]
  end
  AL -->|"EngineCoreRequest<br/>ZMQ + msgpack"| SCH
  EX -->|"SchedulerOutput"| MR
  MR -->|"ModelRunnerOutput<br/>sampled token IDs"| SCH
  SCH -->|"EngineCoreOutputs"| OP
  class H,R,IP io
  class AL,SCH,EX queue
  class OP neutral
  class KV memory
  class MR,M compute

How it really works#

The route, step by step#

Follow one chat request. Class names are real; file paths are relative to the repository root.

#WhereWhat happensCode
1API serverFastAPI receives POST /v1/chat/completions and validates the JSON.vllm/entrypoints/openai/chat_completion/
2API serverThe renderer applies the model’s chat template to the messages and tokenises the result into token IDs. Images and audio are preprocessed here too.vllm/renderers/
3API serverInputProcessor checks the prompt fits max_model_len, fills in sampling defaults and builds an EngineCoreRequest.vllm/v1/engine/input_processor.py
4API serverAsyncLLM.add_request checks admission limits, registers the request with the OutputProcessor, and sends it to the engine core.vllm/v1/engine/async_llm.py
5SocketThe request is serialised with msgpack and sent over a ZeroMQ socket.vllm/v1/engine/core_client.py
6Engine coreAn input thread decodes it, computes the block hashes of its prompt for prefix caching, and puts a Request on the input queue.vllm/v1/engine/core.py
7Engine coreThe busy loop calls scheduler.schedule(). It decides how many tokens of each request run this step and asks the KVCacheManager for blocks. Result: a SchedulerOutput.vllm/v1/core/sched/scheduler.py
8Engine core → workerThe executor sends the SchedulerOutput to every worker of this engine.vllm/v1/executor/multiproc_executor.py
9WorkerThe model runner turns it into tensors — token IDs, positions, block tables — runs the forward pass and samples one token per request that is ready for one.vllm/v1/worker/gpu_model_runner.py
10Engine corescheduler.update_from_output() appends the new tokens, checks stop conditions and frees the blocks of finished requests.vllm/v1/core/sched/scheduler.py
11SocketEngineCoreOutputs — new token IDs for every request that produced any — go back over a second socket.vllm/v1/engine/core.py
12API serverOutputProcessor detokenises incrementally, checks stop strings and puts a RequestOutput on that request’s queue. The HTTP handler formats it as a server-sent event.vllm/v1/engine/output_processor.py

Steps 7 to 11 repeat once per generated token. A 500-token answer goes around that inner loop 500 times; steps 1 to 6 happen once.

The scheduler has no “prefill phase”#

The most useful sentence in the whole codebase is a comment at the top of schedule():

Python
# NOTE(woosuk) on the scheduling algorithm:
# There's no "decoding phase" nor "prefill phase" in the scheduler.
# Each request just has the num_computed_tokens and
# num_tokens_with_spec. num_tokens_with_spec =
# len(prompt_token_ids) + len(output_token_ids) + len(spec_token_ids).
# At each step, the scheduler tries to assign tokens to the requests
# so that each request's num_computed_tokens can catch up its
# num_tokens_with_spec. This is general enough to cover
# chunked prefills, prefix caching, speculative decoding,
# and the "jump decoding" optimization in the future.

Every request carries two counters:

  • how many tokens it has (prompt, plus output so far), and
  • how many of those the GPU has already processed (num_computed_tokens).

Scheduling is closing the gap between them under a budget. A fresh 3,000-token prompt has a gap of 3,000. A request in the middle of generating has a gap of exactly 1 — the token it just sampled. Reading a long prompt in pieces (“chunked prefill”), skipping a prompt that is already cached (“prefix caching”) and checking several guessed tokens at once (“speculative decoding”) are all the same operation with different gap sizes. Hold on to this idea; it makes The Scheduler short.

The codebase in one table#

vllm/ holds a little over a million lines of Python. Almost none of it is on the route above. The sizes are line counts on 5 October 2026.

DirectoryLinesWhat lives there
vllm/v1/178,000The engine. Scheduler, KV cache, engine core, executor, workers, sampling, metrics. Most of this course.
vllm/entrypoints/44,000HTTP servers and CLI: OpenAI, Anthropic and Cohere routes, vllm serve, the offline LLM class
vllm/renderers/8,000Chat templates and tokenisation, per model family
vllm/model_executor/366,000Model implementations in models/, shared layers in layers/ (attention, MoE, quantisation), weight loaders
vllm/models/140,000Newer model families packaged with their own kernels
vllm/distributed/79,000Multi-GPU communication and the KV connectors that move cache between machines
vllm/config/17,000Every configuration object; VllmConfig bundles them all
vllm/compilation/17,000torch.compile integration and CUDA graph wrappers
vllm/tool_parsers/, vllm/reasoning/20,000Turning model text into tool calls and separating reasoning from answers
vllm/lora/, vllm/multimodal/27,000Adapter management; image, audio and video preprocessing
vllm/platforms/, vllm/plugins/6,000Hardware abstraction (CUDA, ROCm, XPU, CPU) and extension points
csrc/—C++ and CUDA: custom kernels, all-reduce, quantisation
rust/—An experimental replacement for the API server, written in Rust

Inside vllm/v1/, these files are the ones to know by name:

vllm/v1/
  engine/
    async_llm.py          the API server's handle on the engine
    core_client.py        the socket client (one class per topology)
    core.py               EngineCore: the busy loop and step()
    input_processor.py    request validation
    output_processor.py   detokenise, stop strings, per-request queues
  core/
    sched/scheduler.py    schedule() and update_from_output()
    kv_cache_manager.py   allocate_slots(), get_computed_blocks()
    block_pool.py         the blocks, the free queue, the hash table
    kv_cache_utils.py     block hashing; turning memory into a block count
  executor/               one engine core talking to N workers
  worker/
    gpu_worker.py         process setup, memory profiling
    gpu_model_runner.py   build the batch, run the model, sample
  attention/backends/     FlashAttention, FlashInfer, Triton, MLA variants
  sample/                 the sampler and logits processors
  structured_output/      grammar-constrained decoding
  spec_decode/            speculative decoding proposers
  metrics/                Prometheus and logging

One object holds every setting#

Every class in the engine is constructed with a single VllmConfig. It bundles the model config, cache config, scheduler config, parallel config and a dozen more. The design choice is deliberate: a new feature that touches only the model runner adds one field to the config instead of threading a new argument through the engine, executor, worker and runner constructors. When this course says “the scheduler reads max_num_batched_tokens”, it reads it from vllm_config.scheduler_config.

“V1” and the names you will meet#

The engine was rewritten in 2025. The rewrite lives in vllm/v1/ and is the only engine today; the old one is gone, but the directory name stayed. Similarly:

NameMeaning
V1The current engine architecture, as opposed to the original 2023 design
Model runner V2A newer rewrite of just the model runner, in vllm/v1/worker/gpu/; both runners exist today
AsyncLLMThe engine as seen from an asyncio program; what the server uses
LLMEngineThe same, synchronous; what the offline LLM class uses
EngineCoreThe scheduler-plus-executor loop itself
APCAutomatic prefix caching
TP / PP / DP / EPTensor, pipeline, data and expert parallelism
P/DPrefill/decode disaggregation: different machines for the two halves of a request

Code#

The whole route, shrunk to a Go program you can hold in your head. There is no model; “running the GPU” just advances counters. What it keeps is the real control flow: a per-step token budget, running requests served before waiting ones, long prompts cut into chunks, and a token sampled only when a request has caught up.

Go
package main

import (
	"fmt"
	"strings"
)

type request struct {
	id       string
	prompt   int // prompt tokens
	maxNew   int // tokens to generate
	output   int // tokens generated so far
	computed int // tokens whose KV cache is on the GPU
}

// total is what vLLM calls num_tokens: prompt plus output so far.
func (r *request) total() int { return r.prompt + r.output }

type slot struct {
	r *request
	n int
}

func main() {
	const budget = 8 // max_num_batched_tokens: tokens the GPU processes per step

	waiting := []*request{
		{id: "A", prompt: 5, maxNew: 3},
		{id: "B", prompt: 12, maxNew: 2},
		{id: "C", prompt: 3, maxNew: 4},
	}
	var running []*request

	for step := 1; len(waiting)+len(running) > 0; step++ {
		left := budget
		var plan []slot

		// 1. Running requests are scheduled first.
		for _, r := range running {
			if left == 0 {
				break
			}
			n := min(r.total()-r.computed, left)
			plan = append(plan, slot{r, n})
			left -= n
		}
		// 2. Waiting requests are admitted while budget remains.
		for len(waiting) > 0 && left > 0 {
			r := waiting[0]
			waiting = waiting[1:]
			running = append(running, r)
			n := min(r.total()-r.computed, left)
			plan = append(plan, slot{r, n})
			left -= n
		}

		// 3. "Forward pass": every scheduled token gets its KV computed.
		var parts, events []string
		for _, s := range plan {
			s.r.computed += s.n
			parts = append(parts, fmt.Sprintf("%s:%d", s.r.id, s.n))
			// 4. A request that has caught up gets one sampled token.
			if s.r.computed == s.r.total() {
				s.r.output++
				events = append(events, fmt.Sprintf("%s->token %d", s.r.id, s.r.output))
			}
		}

		// 5. Finished requests leave and free their memory.
		keep := running[:0]
		for _, r := range running {
			if r.output == r.maxNew {
				events = append(events, r.id+" done")
				continue
			}
			keep = append(keep, r)
		}
		running = keep

		fmt.Printf("step %d  batch [%-11s] used %d/%d  %s\n",
			step, strings.Join(parts, " "), budget-left, budget, strings.Join(events, ", "))
	}
}

Read the output line by line. In step 1, request B gets only 3 of its 12 prompt tokens: that is chunked prefill. In step 2, A is already decoding (one token per step) while B is still reading its prompt in the same batch: that is continuous batching. C waits two steps because the budget is full, and joins in step 3 without anything pausing for it.

Remember this#

  • Three processes: API server, engine core, one worker per GPU.
  • Text becomes tokens once; the schedule-run-sample loop then repeats once per output token.
  • The scheduler has no phases. It closes the gap between tokens a request has and tokens the GPU has computed.
  • schedule() produces a SchedulerOutput; the worker returns a ModelRunnerOutput; update_from_output() closes the step.
  • The engine is vllm/v1/. Model code is five times larger and is not on the control path.
  • One VllmConfig object carries every setting to every class.

Try it#

  1. Raise budget to 32 and rerun. How many steps does B’s prompt take now? What did that do to A’s time between tokens in the early steps? This is the real trade-off behind --max-num-batched-tokens.
  2. Add a fourth request with prompt: 40. Watch it take several steps to read its prompt while the others keep producing tokens.
  3. In a checkout of vLLM, open vllm/v1/engine/core.py and find the method step. Match its five statements to steps 7 to 10 in the table above.

Check yourself#

  1. Which steps of the route happen once per request, and which happen once per token?
  2. A request in the middle of generating text has what gap between its two counters?
  3. In which process does tokenisation happen, and why is it not the engine core?

Sources#

Checked on 5 October 2026 against main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom