Pidoku
The Scheduler

The Scheduler

IntermediateTopic4 lessons~3h 35m

Lessons, in order

About this topic

Once per step, one function decides which tokens the GPU computes next. Its inputs are a token budget, a pool of memory blocks and two queues; its output is a batch. Every latency and throughput property of a vLLM server — time to first token, smoothness of streaming, behaviour under overload — is a consequence of that function and four or five numbers it reads. This topic reads it line by line.

#LessonThe question it answers
01One Scheduling StepWhat exactly do schedule() and update_from_output() do?
02Chunked Prefill and the Token BudgetWhat is max_num_batched_tokens on my GPU, and how do I choose it?
03Priorities, Preemption and QueuesWho waits, who is evicted when memory runs out, and what does eviction cost?
04Async SchedulingHow can a step be planned before the previous one has produced its token?

The general ideas are taught without reference to any engine in Continuous Batching, Chunked Prefill and Scheduling and Priorities. Read those for the why; read this topic for what vLLM actually does.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom