Once per step, one function decides which tokens the GPU computes next. Its inputs are a token budget, a pool of memory blocks and two queues; its output is a batch. Every latency and throughput property of a vLLM server — time to first token, smoothness of streaming, behaviour under overload — is a consequence of that function and four or five numbers it reads. This topic reads it line by line.
| # | Lesson | The question it answers |
|---|---|---|
| 01 | One Scheduling Step | What exactly do schedule() and update_from_output() do? |
| 02 | Chunked Prefill and the Token Budget | What is max_num_batched_tokens on my GPU, and how do I choose it? |
| 03 | Priorities, Preemption and Queues | Who waits, who is evicted when memory runs out, and what does eviction cost? |
| 04 | Async Scheduling | How can a step be planned before the previous one has produced its token? |
The general ideas are taught without reference to any engine in Continuous Batching, Chunked Prefill and Scheduling and Priorities. Read those for the why; read this topic for what vLLM actually does.