Pidoku
KV Cache Management

KV Cache Management

IntermediateTopic5 lessons~4h 10m

Lessons, in order

About this topic

The KV cache is where a serving engine’s memory goes, and how it is managed decides how many requests fit on a GPU. vLLM allocates all of it once, cuts it into blocks, and from then on only moves block numbers between a free queue, the requests that own them and a table of hashes. This topic reads that machinery: the pool, the cache built on top of it, the single function through which every allocation passes, and the arithmetic that sizes it.

#LessonThe question it answers
01Blocks and the PoolWhat is a block, where does a freed block go, and why is that order an eviction policy?
02Prefix CachingHow does vLLM recognise a prompt it has seen before, and why is my hit rate low?
03Allocating Slots: One Call, TracedWhat exactly is checked when the scheduler asks for memory?
04Sizing the Cache: From Gigabytes to BlocksHow many tokens fit on my GPU, and which flags change that?
05More Than One Kind of CacheWhat changes for models whose layers do not all need a full-length cache?

For the ideas without the implementation, see KV Cache, KV Cache Math, Paged Attention and Prefix Caching.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom