Most people run vLLM as a black box with forty flags. This course opens the box. It follows one request through every process, class and data structure of the engine, quotes the real source for each step, and rebuilds the important mechanisms as small Go programs you can run. By the end, every flag is a constant in a loop you have read.
vLLM is the most widely used open-source engine for serving language models. It is also a million lines of Python, C++ and Rust that changes every two weeks, and its documentation explains how to use it far better than how it works. This course is about how it works: why a request waits, where its memory lives, what one engine step does, and which line of code each number in the startup log comes from.
It starts at “what is an inference engine?” and ends with running a fleet on Kubernetes and extending the engine. It assumes you can program in Go and use a terminal. It does not assume you know Python well, have read any vLLM source, or have a GPU.
What you will be able to draw#
flowchart LR
C[":i-users: <b>Clients</b>"] --> API[":vllm: <b>API server</b><br/><small>template, tokenise,<br/>admit, detokenise</small>"]
API -->|"token IDs"| SCH[":i-list-checks: <b>Scheduler</b><br/><small>token budget,<br/>running + waiting</small>"]
SCH --- KV[(":i-layers: <b>Block pool</b><br/><small>paged KV cache,<br/>prefix hashes</small>")]
SCH -->|"SchedulerOutput"| MR[":pytorch: <b>Model runner</b><br/><small>flat batch,<br/>index tensors</small>"]
MR --> ATT[":nvidia: <b>Attention backend</b><br/><small>reads the cache<br/>through a block table</small>"]
MR --> SMP[":i-zap: <b>Sampler</b>"]
SMP -->|"token IDs"| SCH
SCH -->|"new tokens"| API
API -->|"streamed text"| C
class C neutral
class API io
class SCH queue
class KV memory
class MR,ATT,SMP computeEvery box is a topic. Every arrow is a lesson.
How this course works#
Every lesson has the same shape:
- The idea in one minute — the whole lesson in a paragraph.
- A picture — the mechanism, drawn. Press Expand on any diagram to open it full size.
- How it really works — the precise version, with the real class names, the real code and the reasons the authors gave in their own comments.
- Code — a small Go program, standard library only, that reproduces the mechanism or the arithmetic. Each one runs offline and its output is discussed in the text.
- Remember this, Try it and Check yourself.
vLLM itself is written in Python, so quoted source is Python. Everything you are asked to write or run is Go.
The topics#
flowchart LR A["What vLLM is"] --> B["The request path"] B --> C["The scheduler"] C --> D["KV cache"] D --> E["Model execution"] E --> F["Features"] F --> G["Scaling and operating"] G --> H["Extending vLLM"] class A neutral class B io class C queue class D memory class E,F compute class G,H neutral
| Topic | You will be able to | Level |
|---|---|---|
| What vLLM Is | Say what problem paging the KV cache solved, start a server, read its startup log, and place any class on a map | Foundations |
| The Request Path | Follow a request across three processes and two kinds of wire, and say where its latency went | Basic |
| The Scheduler | Read schedule() line by line; choose a token budget; explain preemption and async scheduling | Intermediate |
| KV Cache Management | Explain the block pool and prefix cache from the source, and compute cache capacity for any model and GPU | Intermediate |
| Model Execution | Describe a step inside a worker: the flat batch, attention backends, CUDA graphs, sampling | Advanced |
| Features | Predict the cost of speculative decoding, LoRA, images, tool calling and quantization from where each hooks in | Advanced |
| Scaling and Operating | Lay a model out across GPUs, read the metrics, benchmark honestly and deploy on Kubernetes | Expert |
| Extending vLLM | Choose the right extension point, add a model, and keep your knowledge current | Expert |
34 lessons, 35 Go programs. The glossary defines every term in one line.
How it connects to the other courses#
This course is one engine, in depth. The ideas it implements are taught without reference to any engine elsewhere on this site:
- KV caches, continuous batching, paged attention and prefix caching in general: Inference Engineering. Its lesson vLLM Architecture is a one-page summary of what this course covers in thirty-four.
- What the GPU under the engine is doing: GPU Engineering.
- Instrumenting and watching it: Observability Engineering.
- Where an engine sits in a complete system: AI System Design.
- The Go used in the examples: Golang Engineering.
What you need#
- Go 1.22 or newer and a terminal. Every program runs offline.
- No GPU. A few “Try it” exercises suggest experiments on a real server; they are optional and marked as such.
A promise about versions#
vLLM releases roughly every two weeks, so a course about its internals needs to say exactly what it describes.
- Everything here was read from the vLLM
mainbranch at commit0c16eeeon 5 October 2026, and cross-checked against the v0.30.0 release of 22 September 2026. - Where the two differ, the lesson says which one a statement applies to.
- Every lesson links to the exact source files at that commit, so you can compare with today’s code in one click.
- Where vLLM’s own documentation is out of date with its code — it is in a few places, and it says so itself — the lesson follows the code and notes the difference.
Mechanisms last; defaults do not. The last lesson, The Rust Frontend, and Keeping Up, sorts what you have learned into those two piles and shows how to re-check anything in a few minutes.