Pidoku
Model Execution

Model Execution

AdvancedTopic5 lessons~4h

Lessons, in order

About this topic

The scheduler decides what runs. This topic is about how it runs: the path from a SchedulerOutput to sampled token IDs inside a worker process. It covers the classes that own a GPU, the index tensors that describe an irregular batch, the kernels that read a paged cache, the recordings that remove Python from the hot path, and the pipeline that turns scores into a token.

#LessonThe question it answers
01Executor, Worker, Model RunnerWhich class does what inside a worker, and why are there two model runners?
02Building the BatchHow are tokens from many requests of different lengths fed to the model without padding?
03Attention BackendsWhich attention kernel am I running, how was it chosen, and what does it need?
04CUDA Graphs and CompilationWhy does startup take a minute, and what does that buy?
05Sampling and Structured OutputIn what order are sampling settings applied, and how is valid JSON guaranteed?

Background, without reference to vLLM: CUDA Graphs in Serving, FlashAttention, Sampling and Compilers.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom