Pidoku

More Than One Kind of Cache

Intermediate 45 min Difficulty 4/5 Lesson 05 of 05

Prerequisites Allocating Slots

The idea in one minute#

So far every layer of the model needed the same thing from the cache: room for every token, forever. Many current models break that assumption. Some layers look only at the most recent thousand tokens; some keep a fixed-size running state instead of per-token keys and values; some share another layer’s cache. Giving all of them a full-length cache wastes most of the memory. vLLM’s answer is to split layers into KV cache groups, give each group its own block table per request and its own rule for how many blocks a sequence needs, and draw all of them from the one shared pool. Prefix caching then has to find a prefix that is valid for every group at once.

A picture#

flowchart LR
  REQ[":i-file-text: <b>One request</b><br/><small>112 tokens</small>"] --> CO[":i-split: <b>KVCacheCoordinator</b>"]
  CO --> G0[":i-layers: <b>Group 0 — full attention</b><br/><small>10 layers · needs all 112 tokens<br/>7 blocks</small>"]
  CO --> G1[":i-layers: <b>Group 1 — sliding window</b><br/><small>10 layers · last 32 tokens<br/>2 blocks</small>"]
  CO --> G2[":i-layers: <b>Group 2 — sliding window</b><br/><small>10 layers · last 32 tokens<br/>2 blocks</small>"]
  G0 --> POOL[(":i-database: <b>One BlockPool</b><br/><small>one page size for everyone</small>")]
  G1 --> POOL
  G2 --> POOL
  class REQ neutral
  class CO queue
  class G0,G1,G2 compute
  class POOL memory

How it really works#

Layers that need less#

Layer typeWhat it storesBlocks needed for n tokensExample models
Full attentionKeys and values for every tokenceil(n / block)Llama, Qwen, Mistral
Sliding-window attentionKeys and values for the last w tokens onlyabout ceil(w / block), constantGemma 2 and 3, gpt-oss, Ministral
Chunked local attentionThe current chunk onlyUp to one chunk’s worthLlama 4
State-space (Mamba)A fixed-size state, not per-token dataOne, plus checkpointsJamba, Bamba
Latent attention (MLA)One compressed vector per token instead of separate K and V per headceil(n / block), smaller pagesDeepSeek V3-class
Cross-attentionKeys and values of the encoder’s outputFixed by the encoder lengthWhisper

Each attention layer in a model reports a spec saying which of these it is (FullAttentionSpec, SlidingWindowSpec, MambaSpec, MLAAttentionSpec, … in vllm/v1/kv_cache_interface.py). At startup the engine core collects the specs from the workers and decides how to group them.

Why it matters: the sliding-window case#

Gemma-3-27B has 62 layers: 10 full attention and 52 with a sliding window of 1,024 tokens. For a 32,000-token prompt:

treat all 62 layers as full attention     62 × 32,000         = 1,984,000 layer-tokens
honour the window                         10 × 32,000 + 52 × 1,024 = 373,248 layer-tokens

                                                               5.3 times less memory

That is the difference between fitting three such requests and fitting sixteen.

The constraint: one page size#

The block pool hands out identical blocks. A block is a fixed number of bytes, a page. But a group of 10 layers and a group of 52 layers would need pages of different sizes for the same number of tokens. The design question is how to make every group use the same page size.

The answer is to make groups the same size. From the design document:

  • A regular ratio. With 20 sliding-window and 10 full layers, form three groups of 10. Each group’s page is 10 × block_size × bytes per token per layer.
  • No regular ratio. Gemma-3-27B’s 52 and 10 do not divide evenly. vLLM uses the smaller count as the group size: one full group of 10, five sliding-window groups of 10, and a seventh group with the 2 remaining layers and 8 layers of padding. The document is frank about it: “The solution has some memory waste and is not perfect.”
  • Different sizes per token (state-space layers, whose state is much larger than one token’s keys and values). The attention layers’ block_size is increased until one attention page is at least as large as the state, and the state is padded up to match. This is why some hybrid models run with blocks of several hundred or a thousand tokens rather than 16, and why --prefix-match-unit exists.

A request then has one block table per group. You can see it in the type of NewRequestData.block_ids: tuple[list[int], ...], one list per group.

Layers share buffers across groups#

Physical memory is not one tensor per group. For n groups of m layers each, vLLM allocates m buffers, and each buffer is shared by n layers, one from each group. Layer full.0, layer sw.0 and layer sw.10 all read and write the same tensor, at different block indices. Since the three groups hold disjoint sets of block IDs, they never collide. The benefit is that “block 37” means the same thing for every layer: a slice of the same size at the same offset.

Freeing blocks as the window moves#

A sliding-window layer does not use a ring buffer. A ring would overwrite old tokens in place, and a block whose contents change cannot be shared or cached. Instead vLLM gives each token its own slot, exactly as for full attention, and frees blocks once they fall behind the window:

Python
# Free the blocks that are skipped during the attention computation
# (e.g., tokens outside the sliding window).
self.coordinator.remove_skipped_blocks(request.request_id, ...)

The freed positions in the request’s table are replaced by the null block (block 0), so the table keeps its shape and indices still line up with token positions. The freed blocks return to the pool with their hashes intact and may still serve another request’s cache hit.

Prefix caching across groups#

A cache hit is only usable if every layer can resume from it. The rules differ by type:

TypeA prefix of length L is a valid hit if…
Full attentionevery block from 0 to L is cached
Sliding windowthe blocks covering the last w − 1 tokens before L are cached
State-spacea saved state exists exactly at L

Full attention scans left to right and stops at the first miss. A sliding-window group can have a hit at a position where its earliest blocks are long gone, so it is scanned the other way, right to left, taking the first position that works.

The coordinator combines them:

  1. Find the longest hit for the full-attention group, scanning from the left.
  2. Within that length, find the longest hit for the other group, scanning from the right.

The result is valid for both. The design document notes the efficiency: the second scan starts from the first scan’s answer, so when there is no hit at all it stops immediately.

The cache key includes the group: internally the map is from (block hash, group id) to a block. The same tokens are cached, and evicted, independently for each group.

Which coordinator you get#

get_kv_cache_coordinator chooses one at startup:

CoordinatorWhen
KVCacheCoordinatorNoPrefixCachePrefix caching is disabled. Any number of groups; no lookups.
UnitaryKVCacheCoordinatorExactly one group. The common case; no intersection needed.
HybridKVCacheCoordinatorSeveral groups with prefix caching on

note

The published design document, written against an older commit, says the hybrid coordinator “handles exactly two KV cache groups”. The class on main has grown since — it now deals with state-space checkpoints, bounded replay for sliding windows and several specialised layouts. Treat the document as the explanation of the idea and the source as the truth about current limits.

Turning it off#

--disable-hybrid-kv-cache-manager makes vLLM treat every attention layer as full attention for allocation, while the model still computes with its real window. It wastes memory — this is the first row of the arithmetic above — but it removes a whole class of complexity, and a few features have required it at one time or another. If the startup log says the hybrid manager was disabled, your capacity is the full-attention figure.

KV sharing#

In some models a layer does not have its own cache at all; it reads another layer’s. vLLM allocates nothing for such layers and points them at the owner’s buffer. They add nothing to the per-token cost, which is why counting layers in config.json can overestimate the cache for these architectures.

What this changes for you#

For Llama-style models: nothing. One group, one table, the arithmetic of Sizing the Cache.

For hybrid models:

  • Capacity is not blocks × 16. A request occupies blocks in several groups, a different number in each. vLLM computes a “group-aware” capacity and that is what the startup log prints as GPU KV cache size: N tokens. Trust that line over hand arithmetic.
  • Per-token cost is not constant. Short requests are relatively expensive (the window is not yet full); long ones are cheap per token.
  • Block size may be large, so prefix-cache hits are coarse unless --prefix-match-unit is set.
  • New architectures arrive with new specs. When a freshly released model is slow or memory-hungry in vLLM, the usual cause is that its layers are being treated as full attention until a dedicated spec lands.

Code#

Group formation and per-request memory for hybrid models, including the padding waste.

Go
package main

import "fmt"

const blockSize = 16

func ceilDiv(a, b int) int { return (a + b - 1) / b }

// blocksFull: every token keeps its slot.
func blocksFull(n int) int { return ceilDiv(n, blockSize) }

// blocksWindow: slots for the last w tokens; whole blocks behind the window are freed.
func blocksWindow(n, w int) int {
	skipped := n - w
	if skipped < 0 {
		skipped = 0
	}
	return ceilDiv(n, blockSize) - skipped/blockSize
}

type model struct {
	name          string
	fullLayers    int
	windowLayers  int
	window        int
	groupSize     int // layers per group = the smaller layer count
	windowGroups  int
	paddingLayers int
}

func newModel(name string, full, win, window int) model {
	m := model{name: name, fullLayers: full, windowLayers: win, window: window}
	m.groupSize = min(full, win)
	m.windowGroups = ceilDiv(win, m.groupSize)
	m.paddingLayers = m.windowGroups*m.groupSize - win
	return m
}

// cost returns layer-tokens of memory for one request of n tokens:
// blocks x block size x layers per group, summed over groups.
func (m model) cost(n int) (hybrid, allFull int) {
	page := blockSize * m.groupSize
	hybrid = blocksFull(n)*page*(m.fullLayers/m.groupSize) + blocksWindow(n, m.window)*page*m.windowGroups
	allFull = blocksFull(n) * blockSize * (m.fullLayers + m.windowLayers)
	return
}

func main() {
	// The example from vLLM's design document: 10 full + 20 window layers, window 32.
	toy := newModel("toy", 10, 20, 32)
	fmt.Printf("%s: %d groups of %d layers (1 full, %d window), padding %d layers\n",
		toy.name, 1+toy.windowGroups, toy.groupSize, toy.windowGroups, toy.paddingLayers)
	fmt.Printf("  112-token request: %d blocks for the full group, %d for each window group, %d in total\n\n",
		blocksFull(112), blocksWindow(112, 32), blocksFull(112)+toy.windowGroups*blocksWindow(112, 32))

	g := newModel("Gemma-3-27B", 10, 52, 1024)
	fmt.Printf("%s: %d full + %d window layers -> 1 full group + %d window groups of %d, padding %d layers\n\n",
		g.name, g.fullLayers, g.windowLayers, g.windowGroups, g.groupSize, g.paddingLayers)

	fmt.Printf("%10s %16s %16s %10s\n", "tokens", "hybrid manager", "all-full", "saving")
	for _, n := range []int{512, 2000, 8000, 32000, 128000} {
		h, f := g.cost(n)
		fmt.Printf("%10d %16d %16d %9.1fx\n", n, h, f, float64(f)/float64(h))
	}
	fmt.Println("\n(units: layer-tokens; multiply by bytes per token per layer for bytes)")
}

Two things to read from the output. For a 512-token request the hybrid manager is worse than treating every layer as full attention: the window is not full yet, so nothing is saved, and the eight padding layers cost extra. The saving appears once requests exceed the window and grows with length. A model like this is a good fit for long-context work and a slightly poor one for short chats — which is exactly what its architecture was designed for.

Remember this#

  • Layers that need different amounts of cache are placed in different KV cache groups.
  • All groups draw equal-sized pages from one pool, so groups are made the same size, with padding if needed.
  • A request has one block table per group.
  • Sliding-window layers free blocks behind the window and leave null blocks in their place.
  • A prefix-cache hit must be valid for every group: left-to-right for full attention, right-to-left for the rest.
  • For hybrid models, read capacity from the startup log; blocks × 16 is wrong.
  • --disable-hybrid-kv-cache-manager trades memory for simplicity.

Try it#

  1. Add a model with 30 full and 30 window layers and a window of 4,096. How many groups, how much padding, and what saving at 32,000 tokens?
  2. For Gemma-3-27B, find the request length at which the hybrid manager breaks even with all-full. Why is it a little above the window size and not exactly at it?
  3. Change groupSize to the greatest common divisor of the two layer counts instead of the minimum. For 52 and 10 that is 2. How many groups result, and what did vLLM gain by accepting padding instead?

Check yourself#

  1. Why must all KV cache groups have the same page size?
  2. Why does a sliding-window layer free old blocks instead of overwriting them in a ring?
  3. For a model with full and sliding-window layers, in which direction is each group scanned for a cache hit, and why?

Sources#

Checked on 5 October 2026 against main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom