Pidoku
What vLLM Is

What vLLM Is

FoundationsTopic3 lessons~1h 55m

Lessons, in order

About this topic

Start here. These three lessons need no background beyond a terminal. They explain the problem vLLM was built to solve, get a server running, and walk one request through the whole system so that every later lesson has a place on the map.

#LessonThe question it answers
01What vLLM Is and Why It ExistsWhat does an inference engine do, and what was vLLM’s founding idea?
02Your First Server, and How to Read Its Startup LogHow do I run it, and what do the numbers it prints at startup mean?
03The Map: One Request, End to EndWhich processes and classes does a request pass through, and where is each in the source?

If the words token, KV cache or prefill are new, read Transformer Inference Overview and KV Cache alongside lesson 01.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom