The KV cache is where a serving engine’s memory goes, and how it is managed decides how many requests fit on a GPU. vLLM allocates all of it once, cuts it into blocks, and from then on only moves block numbers between a free queue, the requests that own them and a table of hashes. This topic reads that machinery: the pool, the cache built on top of it, the single function through which every allocation passes, and the arithmetic that sizes it.
| # | Lesson | The question it answers |
|---|---|---|
| 01 | Blocks and the Pool | What is a block, where does a freed block go, and why is that order an eviction policy? |
| 02 | Prefix Caching | How does vLLM recognise a prompt it has seen before, and why is my hit rate low? |
| 03 | Allocating Slots: One Call, Traced | What exactly is checked when the scheduler asks for memory? |
| 04 | Sizing the Cache: From Gigabytes to Blocks | How many tokens fit on my GPU, and which flags change that? |
| 05 | More Than One Kind of Cache | What changes for models whose layers do not all need a full-length cache? |
For the ideas without the implementation, see KV Cache, KV Cache Math, Paged Attention and Prefix Caching.