Pidoku

vLLM

vLLM from the inside — the processes, the scheduler, the KV cache block pool, prefix caching, the model runner and the kernels, read from the source and explained line by line.

Start reading
  • 5 levels
  • 8 topics
  • 34 lessons
  • ~26h 20m

Contents

  1. Foundations

    Build the mental model.

    3 lessons · ~1h 55m

    1. What vLLM IsStart here. These three lessons need no background beyond a terminal. They explain the problem vLLM was built to solve, get a server running, and walk one request through the whole system so … 3 lessons · ~1h 55m
      1. ·Overview
      2. 01What vLLM Is and Why It ExistsFoundations30 min
      3. 02Your First Server, and How to Read Its Startup LogFoundations40 min
      4. 03The Map: One Request, End to EndFoundations45 min
  2. Basic

    Understand the core mechanisms.

    4 lessons · ~3h 15m

    1. The Request PathA request crosses three processes on its way to the GPU and back. This topic follows it across each boundary: what is sent, in what format, by which thread, and what can go wrong there. By … 4 lessons · ~3h 15m
      1. ·Overview
      2. 01Processes and WiresBasic50 min
      3. 02The Frontend: From JSON to Token IDsBasic45 min
      4. 03AsyncLLM and the Way Back: Tokens to TextBasic50 min
      5. 04The Engine Core LoopBasic50 min
  3. Intermediate

    Learn the optimization techniques.

    9 lessons · ~7h 45m

    1. The SchedulerOnce per step, one function decides which tokens the GPU computes next. Its inputs are a token budget, a pool of memory blocks and two queues; its output is a batch. Every latency and … 4 lessons · ~3h 35m
      1. ·Overview
      2. 01One Scheduling StepIntermediate1h
      3. 02Chunked Prefill and the Token BudgetIntermediate50 min
      4. 03Priorities, Preemption and QueuesIntermediate50 min
      5. 04Async Scheduling: Planning a Step Before the Last One FinishesIntermediate55 min
    2. KV Cache ManagementThe KV cache is where a serving engine's memory goes, and how it is managed decides how many requests fit on a GPU. vLLM allocates all of it once, cuts it into blocks, and from then on only … 5 lessons · ~4h 10m
      1. ·Overview
      2. 01Blocks and the PoolIntermediate55 min
      3. 02Prefix CachingIntermediate55 min
      4. 03Allocating Slots: One Call, TracedIntermediate45 min
      5. 04Sizing the Cache: From Gigabytes to BlocksIntermediate50 min
      6. 05More Than One Kind of CacheIntermediate45 min
  4. Advanced

    Study systems at production scale.

    10 lessons · ~7h 35m

    1. Model ExecutionThe scheduler decides what runs. This topic is about how it runs: the path from a SchedulerOutput to sampled token IDs inside a worker process. It covers the classes that own a GPU, the … 5 lessons · ~4h
      1. ·Overview
      2. 01Executor, Worker, Model RunnerAdvanced45 min
      3. 02Building the BatchAdvanced50 min
      4. 03Attention BackendsAdvanced45 min
      5. 04CUDA Graphs and CompilationAdvanced50 min
      6. 05Sampling and Structured OutputAdvanced50 min
    2. FeaturesEverything in the earlier topics is the same for every model and every request. This topic covers what changes when you turn something on: guessing tokens ahead, serving many fine-tunes from … 5 lessons · ~3h 35m
      1. ·Overview
      2. 01Speculative DecodingAdvanced50 min
      3. 02LoRA Adapters: Many Models in OneAdvanced40 min
      4. 03Multimodal InputsAdvanced45 min
      5. 04Tool Calling and ReasoningAdvanced40 min
      6. 05Quantization in vLLMAdvanced40 min
  5. Expert

    Design platforms and read the frontier.

    8 lessons · ~5h 50m

    1. Scaling and OperatingOne engine on one GPU is the unit. This topic is about everything beyond it: spreading a model over several GPUs or running many copies, moving cached state between machines, reading the … 5 lessons · ~4h
      1. ·Overview
      2. 01Parallelism: TP, PP, DP, EP and CPExpert50 min
      3. 02Disaggregation and KV ConnectorsExpert50 min
      4. 03Metrics: Reading a Running ServerExpert45 min
      5. 04Tuning and BenchmarkingExpert50 min
      6. 05Deploying on KubernetesExpert45 min
    2. Extending vLLMThe last topic turns from reading vLLM to changing it. It covers the registered extension points that make a fork unnecessary, what a model has to look like to run inside the engine, and how … 3 lessons · ~1h 50m
      1. ·Overview
      2. 01Plugins and Extension PointsExpert35 min
      3. 02Adding a ModelExpert40 min
      4. 03The Rust Frontend, and Keeping UpExpert35 min

Reference

About

Most people run vLLM as a black box with forty flags. This course opens the box. It follows one request through every process, class and data structure of the engine, quotes the real source for each step, and rebuilds the important mechanisms as small Go programs you can run. By the end, every flag is a constant in a loop you have read.

vLLM is the most widely used open-source engine for serving language models. It is also a million lines of Python, C++ and Rust that changes every two weeks, and its documentation explains how to use it far better than how it works. This course is about how it works: why a request waits, where its memory lives, what one engine step does, and which line of code each number in the startup log comes from.

It starts at “what is an inference engine?” and ends with running a fleet on Kubernetes and extending the engine. It assumes you can program in Go and use a terminal. It does not assume you know Python well, have read any vLLM source, or have a GPU.

What you will be able to draw#

flowchart LR
  C[":i-users: <b>Clients</b>"] --> API[":vllm: <b>API server</b><br/><small>template, tokenise,<br/>admit, detokenise</small>"]
  API -->|"token IDs"| SCH[":i-list-checks: <b>Scheduler</b><br/><small>token budget,<br/>running + waiting</small>"]
  SCH --- KV[(":i-layers: <b>Block pool</b><br/><small>paged KV cache,<br/>prefix hashes</small>")]
  SCH -->|"SchedulerOutput"| MR[":pytorch: <b>Model runner</b><br/><small>flat batch,<br/>index tensors</small>"]
  MR --> ATT[":nvidia: <b>Attention backend</b><br/><small>reads the cache<br/>through a block table</small>"]
  MR --> SMP[":i-zap: <b>Sampler</b>"]
  SMP -->|"token IDs"| SCH
  SCH -->|"new tokens"| API
  API -->|"streamed text"| C
  class C neutral
  class API io
  class SCH queue
  class KV memory
  class MR,ATT,SMP compute

Every box is a topic. Every arrow is a lesson.

How this course works#

Every lesson has the same shape:

  1. The idea in one minute — the whole lesson in a paragraph.
  2. A picture — the mechanism, drawn. Press Expand on any diagram to open it full size.
  3. How it really works — the precise version, with the real class names, the real code and the reasons the authors gave in their own comments.
  4. Code — a small Go program, standard library only, that reproduces the mechanism or the arithmetic. Each one runs offline and its output is discussed in the text.
  5. Remember this, Try it and Check yourself.

vLLM itself is written in Python, so quoted source is Python. Everything you are asked to write or run is Go.

The topics#

flowchart LR
  A["What vLLM is"] --> B["The request path"]
  B --> C["The scheduler"]
  C --> D["KV cache"]
  D --> E["Model execution"]
  E --> F["Features"]
  F --> G["Scaling and operating"]
  G --> H["Extending vLLM"]
  class A neutral
  class B io
  class C queue
  class D memory
  class E,F compute
  class G,H neutral
TopicYou will be able toLevel
What vLLM IsSay what problem paging the KV cache solved, start a server, read its startup log, and place any class on a mapFoundations
The Request PathFollow a request across three processes and two kinds of wire, and say where its latency wentBasic
The SchedulerRead schedule() line by line; choose a token budget; explain preemption and async schedulingIntermediate
KV Cache ManagementExplain the block pool and prefix cache from the source, and compute cache capacity for any model and GPUIntermediate
Model ExecutionDescribe a step inside a worker: the flat batch, attention backends, CUDA graphs, samplingAdvanced
FeaturesPredict the cost of speculative decoding, LoRA, images, tool calling and quantization from where each hooks inAdvanced
Scaling and OperatingLay a model out across GPUs, read the metrics, benchmark honestly and deploy on KubernetesExpert
Extending vLLMChoose the right extension point, add a model, and keep your knowledge currentExpert

34 lessons, 35 Go programs. The glossary defines every term in one line.

How it connects to the other courses#

This course is one engine, in depth. The ideas it implements are taught without reference to any engine elsewhere on this site:

What you need#

  • Go 1.22 or newer and a terminal. Every program runs offline.
  • No GPU. A few “Try it” exercises suggest experiments on a real server; they are optional and marked as such.

A promise about versions#

vLLM releases roughly every two weeks, so a course about its internals needs to say exactly what it describes.

  • Everything here was read from the vLLM main branch at commit 0c16eee on 5 October 2026, and cross-checked against the v0.30.0 release of 22 September 2026.
  • Where the two differ, the lesson says which one a statement applies to.
  • Every lesson links to the exact source files at that commit, so you can compare with today’s code in one click.
  • Where vLLM’s own documentation is out of date with its code — it is in a few places, and it says so itself — the lesson follows the code and notes the difference.

Mechanisms last; defaults do not. The last lesson, The Rust Frontend, and Keeping Up, sorts what you have learned into those two piles and shows how to re-check anything in a few minutes.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom