The idea in one minute#
A request to vLLM crosses three processes and about a dozen named objects. Text becomes token IDs in the API server. Token IDs cross a socket to the engine core, where a scheduler decides — once per step — how many tokens of each request the GPU will process, and a KV cache manager hands out the memory blocks for them. The schedule crosses a second boundary to the worker, whose model runner builds tensors and calls the model. One new token per request comes back the same way, is turned into text, and is streamed to the client. Then the loop runs again. Everything else in this course is a closer look at one stop on this route.
A picture#
flowchart TB
subgraph P0["Process 1 — API server (asyncio)"]
direction LR
H[":i-globe: <b>HTTP handler</b><br/><small>FastAPI route</small>"] --> R[":i-file-text: <b>Renderer</b><br/><small>chat template, tokenise</small>"]
R --> IP[":i-funnel: <b>InputProcessor</b><br/><small>validate, build request</small>"]
IP --> AL[":vllm: <b>AsyncLLM</b><br/><small>admission, per-request queue</small>"]
OP[":i-message-square: <b>OutputProcessor</b><br/><small>detokenise, stop strings</small>"] --> H
end
subgraph P1["Process 2 — Engine core (busy loop)"]
direction LR
SCH[":i-list-checks: <b>Scheduler</b><br/><small>token budget, queues</small>"] --- KV[(":i-layers: <b>KVCacheManager</b><br/><small>block pool, prefix cache</small>")]
SCH --> EX[":i-split: <b>Executor</b><br/><small>fan out to workers</small>"]
end
subgraph P2["Process 3..N — one worker per GPU"]
direction LR
MR[":pytorch: <b>Model runner</b><br/><small>build batch, sample</small>"] --> M[":nvidia: <b>Model + attention kernel</b>"]
end
AL -->|"EngineCoreRequest<br/>ZMQ + msgpack"| SCH
EX -->|"SchedulerOutput"| MR
MR -->|"ModelRunnerOutput<br/>sampled token IDs"| SCH
SCH -->|"EngineCoreOutputs"| OP
class H,R,IP io
class AL,SCH,EX queue
class OP neutral
class KV memory
class MR,M computeHow it really works#
The route, step by step#
Follow one chat request. Class names are real; file paths are relative to the repository root.
| # | Where | What happens | Code |
|---|---|---|---|
| 1 | API server | FastAPI receives POST /v1/chat/completions and validates the JSON. | vllm/entrypoints/openai/chat_completion/ |
| 2 | API server | The renderer applies the model’s chat template to the messages and tokenises the result into token IDs. Images and audio are preprocessed here too. | vllm/renderers/ |
| 3 | API server | InputProcessor checks the prompt fits max_model_len, fills in sampling defaults and builds an EngineCoreRequest. | vllm/v1/engine/input_processor.py |
| 4 | API server | AsyncLLM.add_request checks admission limits, registers the request with the OutputProcessor, and sends it to the engine core. | vllm/v1/engine/async_llm.py |
| 5 | Socket | The request is serialised with msgpack and sent over a ZeroMQ socket. | vllm/v1/engine/core_client.py |
| 6 | Engine core | An input thread decodes it, computes the block hashes of its prompt for prefix caching, and puts a Request on the input queue. | vllm/v1/engine/core.py |
| 7 | Engine core | The busy loop calls scheduler.schedule(). It decides how many tokens of each request run this step and asks the KVCacheManager for blocks. Result: a SchedulerOutput. | vllm/v1/core/sched/scheduler.py |
| 8 | Engine core → worker | The executor sends the SchedulerOutput to every worker of this engine. | vllm/v1/executor/multiproc_executor.py |
| 9 | Worker | The model runner turns it into tensors — token IDs, positions, block tables — runs the forward pass and samples one token per request that is ready for one. | vllm/v1/worker/gpu_model_runner.py |
| 10 | Engine core | scheduler.update_from_output() appends the new tokens, checks stop conditions and frees the blocks of finished requests. | vllm/v1/core/sched/scheduler.py |
| 11 | Socket | EngineCoreOutputs — new token IDs for every request that produced any — go back over a second socket. | vllm/v1/engine/core.py |
| 12 | API server | OutputProcessor detokenises incrementally, checks stop strings and puts a RequestOutput on that request’s queue. The HTTP handler formats it as a server-sent event. | vllm/v1/engine/output_processor.py |
Steps 7 to 11 repeat once per generated token. A 500-token answer goes around that inner loop 500 times; steps 1 to 6 happen once.
The scheduler has no “prefill phase”#
The most useful sentence in the whole codebase is a comment at the top of schedule():
# NOTE(woosuk) on the scheduling algorithm:
# There's no "decoding phase" nor "prefill phase" in the scheduler.
# Each request just has the num_computed_tokens and
# num_tokens_with_spec. num_tokens_with_spec =
# len(prompt_token_ids) + len(output_token_ids) + len(spec_token_ids).
# At each step, the scheduler tries to assign tokens to the requests
# so that each request's num_computed_tokens can catch up its
# num_tokens_with_spec. This is general enough to cover
# chunked prefills, prefix caching, speculative decoding,
# and the "jump decoding" optimization in the future.Every request carries two counters:
- how many tokens it has (prompt, plus output so far), and
- how many of those the GPU has already processed (
num_computed_tokens).
Scheduling is closing the gap between them under a budget. A fresh 3,000-token prompt has a gap of 3,000. A request in the middle of generating has a gap of exactly 1 — the token it just sampled. Reading a long prompt in pieces (“chunked prefill”), skipping a prompt that is already cached (“prefix caching”) and checking several guessed tokens at once (“speculative decoding”) are all the same operation with different gap sizes. Hold on to this idea; it makes The Scheduler short.
The codebase in one table#
vllm/ holds a little over a million lines of Python. Almost none of it is on the route above.
The sizes are line counts on 5 October 2026.
| Directory | Lines | What lives there |
|---|---|---|
vllm/v1/ | 178,000 | The engine. Scheduler, KV cache, engine core, executor, workers, sampling, metrics. Most of this course. |
vllm/entrypoints/ | 44,000 | HTTP servers and CLI: OpenAI, Anthropic and Cohere routes, vllm serve, the offline LLM class |
vllm/renderers/ | 8,000 | Chat templates and tokenisation, per model family |
vllm/model_executor/ | 366,000 | Model implementations in models/, shared layers in layers/ (attention, MoE, quantisation), weight loaders |
vllm/models/ | 140,000 | Newer model families packaged with their own kernels |
vllm/distributed/ | 79,000 | Multi-GPU communication and the KV connectors that move cache between machines |
vllm/config/ | 17,000 | Every configuration object; VllmConfig bundles them all |
vllm/compilation/ | 17,000 | torch.compile integration and CUDA graph wrappers |
vllm/tool_parsers/, vllm/reasoning/ | 20,000 | Turning model text into tool calls and separating reasoning from answers |
vllm/lora/, vllm/multimodal/ | 27,000 | Adapter management; image, audio and video preprocessing |
vllm/platforms/, vllm/plugins/ | 6,000 | Hardware abstraction (CUDA, ROCm, XPU, CPU) and extension points |
csrc/ | — | C++ and CUDA: custom kernels, all-reduce, quantisation |
rust/ | — | An experimental replacement for the API server, written in Rust |
Inside vllm/v1/, these files are the ones to know by name:
vllm/v1/
engine/
async_llm.py the API server's handle on the engine
core_client.py the socket client (one class per topology)
core.py EngineCore: the busy loop and step()
input_processor.py request validation
output_processor.py detokenise, stop strings, per-request queues
core/
sched/scheduler.py schedule() and update_from_output()
kv_cache_manager.py allocate_slots(), get_computed_blocks()
block_pool.py the blocks, the free queue, the hash table
kv_cache_utils.py block hashing; turning memory into a block count
executor/ one engine core talking to N workers
worker/
gpu_worker.py process setup, memory profiling
gpu_model_runner.py build the batch, run the model, sample
attention/backends/ FlashAttention, FlashInfer, Triton, MLA variants
sample/ the sampler and logits processors
structured_output/ grammar-constrained decoding
spec_decode/ speculative decoding proposers
metrics/ Prometheus and loggingOne object holds every setting#
Every class in the engine is constructed with a single VllmConfig. It bundles the model
config, cache config, scheduler config, parallel config and a dozen more. The design choice
is deliberate: a new feature that touches only the model runner adds one field to the config
instead of threading a new argument through the engine, executor, worker and runner
constructors. When this course says “the scheduler reads max_num_batched_tokens”, it reads
it from vllm_config.scheduler_config.
“V1” and the names you will meet#
The engine was rewritten in 2025. The rewrite lives in vllm/v1/ and is the only engine
today; the old one is gone, but the directory name stayed. Similarly:
| Name | Meaning |
|---|---|
| V1 | The current engine architecture, as opposed to the original 2023 design |
| Model runner V2 | A newer rewrite of just the model runner, in vllm/v1/worker/gpu/; both runners exist today |
AsyncLLM | The engine as seen from an asyncio program; what the server uses |
LLMEngine | The same, synchronous; what the offline LLM class uses |
EngineCore | The scheduler-plus-executor loop itself |
| APC | Automatic prefix caching |
| TP / PP / DP / EP | Tensor, pipeline, data and expert parallelism |
| P/D | Prefill/decode disaggregation: different machines for the two halves of a request |
Code#
The whole route, shrunk to a Go program you can hold in your head. There is no model; “running the GPU” just advances counters. What it keeps is the real control flow: a per-step token budget, running requests served before waiting ones, long prompts cut into chunks, and a token sampled only when a request has caught up.
package main
import (
"fmt"
"strings"
)
type request struct {
id string
prompt int // prompt tokens
maxNew int // tokens to generate
output int // tokens generated so far
computed int // tokens whose KV cache is on the GPU
}
// total is what vLLM calls num_tokens: prompt plus output so far.
func (r *request) total() int { return r.prompt + r.output }
type slot struct {
r *request
n int
}
func main() {
const budget = 8 // max_num_batched_tokens: tokens the GPU processes per step
waiting := []*request{
{id: "A", prompt: 5, maxNew: 3},
{id: "B", prompt: 12, maxNew: 2},
{id: "C", prompt: 3, maxNew: 4},
}
var running []*request
for step := 1; len(waiting)+len(running) > 0; step++ {
left := budget
var plan []slot
// 1. Running requests are scheduled first.
for _, r := range running {
if left == 0 {
break
}
n := min(r.total()-r.computed, left)
plan = append(plan, slot{r, n})
left -= n
}
// 2. Waiting requests are admitted while budget remains.
for len(waiting) > 0 && left > 0 {
r := waiting[0]
waiting = waiting[1:]
running = append(running, r)
n := min(r.total()-r.computed, left)
plan = append(plan, slot{r, n})
left -= n
}
// 3. "Forward pass": every scheduled token gets its KV computed.
var parts, events []string
for _, s := range plan {
s.r.computed += s.n
parts = append(parts, fmt.Sprintf("%s:%d", s.r.id, s.n))
// 4. A request that has caught up gets one sampled token.
if s.r.computed == s.r.total() {
s.r.output++
events = append(events, fmt.Sprintf("%s->token %d", s.r.id, s.r.output))
}
}
// 5. Finished requests leave and free their memory.
keep := running[:0]
for _, r := range running {
if r.output == r.maxNew {
events = append(events, r.id+" done")
continue
}
keep = append(keep, r)
}
running = keep
fmt.Printf("step %d batch [%-11s] used %d/%d %s\n",
step, strings.Join(parts, " "), budget-left, budget, strings.Join(events, ", "))
}
}Read the output line by line. In step 1, request B gets only 3 of its 12 prompt tokens: that is chunked prefill. In step 2, A is already decoding (one token per step) while B is still reading its prompt in the same batch: that is continuous batching. C waits two steps because the budget is full, and joins in step 3 without anything pausing for it.
Remember this#
- Three processes: API server, engine core, one worker per GPU.
- Text becomes tokens once; the schedule-run-sample loop then repeats once per output token.
- The scheduler has no phases. It closes the gap between tokens a request has and tokens the GPU has computed.
schedule()produces aSchedulerOutput; the worker returns aModelRunnerOutput;update_from_output()closes the step.- The engine is
vllm/v1/. Model code is five times larger and is not on the control path. - One
VllmConfigobject carries every setting to every class.
Try it#
- Raise
budgetto 32 and rerun. How many steps does B’s prompt take now? What did that do to A’s time between tokens in the early steps? This is the real trade-off behind--max-num-batched-tokens. - Add a fourth request with
prompt: 40. Watch it take several steps to read its prompt while the others keep producing tokens. - In a checkout of vLLM, open
vllm/v1/engine/core.pyand find the methodstep. Match its five statements to steps 7 to 10 in the table above.
Check yourself#
- Which steps of the route happen once per request, and which happen once per token?
- A request in the middle of generating text has what gap between its two counters?
- In which process does tokenisation happen, and why is it not the engine core?
Sources#
Checked on 5 October 2026 against main at commit 0c16eee.
vllm/v1/core/sched/scheduler.py—schedule()and the comment quoted abovevllm/v1/engine/core.py—EngineCore.step()vllm/v1/engine/async_llm.py- Architecture overview