Start here. These three lessons need no background beyond a terminal. They explain the problem vLLM was built to solve, get a server running, and walk one request through the whole system so that every later lesson has a place on the map.
| # | Lesson | The question it answers |
|---|---|---|
| 01 | What vLLM Is and Why It Exists | What does an inference engine do, and what was vLLM’s founding idea? |
| 02 | Your First Server, and How to Read Its Startup Log | How do I run it, and what do the numbers it prints at startup mean? |
| 03 | The Map: One Request, End to End | Which processes and classes does a request pass through, and where is each in the source? |
If the words token, KV cache or prefill are new, read Transformer Inference Overview and KV Cache alongside lesson 01.