The scheduler decides what runs. This topic is about how it runs: the path from a
SchedulerOutput to sampled token IDs inside a worker process. It covers the classes that own
a GPU, the index tensors that describe an irregular batch, the kernels that read a paged cache,
the recordings that remove Python from the hot path, and the pipeline that turns scores into a
token.
| # | Lesson | The question it answers |
|---|---|---|
| 01 | Executor, Worker, Model Runner | Which class does what inside a worker, and why are there two model runners? |
| 02 | Building the Batch | How are tokens from many requests of different lengths fed to the model without padding? |
| 03 | Attention Backends | Which attention kernel am I running, how was it chosen, and what does it need? |
| 04 | CUDA Graphs and Compilation | Why does startup take a minute, and what does that buy? |
| 05 | Sampling and Structured Output | In what order are sampling settings applied, and how is valid JSON guaranteed? |
Background, without reference to vLLM: CUDA Graphs in Serving, FlashAttention, Sampling and Compilers.