The idea in one minute#
So far every layer of the model needed the same thing from the cache: room for every token, forever. Many current models break that assumption. Some layers look only at the most recent thousand tokens; some keep a fixed-size running state instead of per-token keys and values; some share another layer’s cache. Giving all of them a full-length cache wastes most of the memory. vLLM’s answer is to split layers into KV cache groups, give each group its own block table per request and its own rule for how many blocks a sequence needs, and draw all of them from the one shared pool. Prefix caching then has to find a prefix that is valid for every group at once.
A picture#
flowchart LR
REQ[":i-file-text: <b>One request</b><br/><small>112 tokens</small>"] --> CO[":i-split: <b>KVCacheCoordinator</b>"]
CO --> G0[":i-layers: <b>Group 0 — full attention</b><br/><small>10 layers · needs all 112 tokens<br/>7 blocks</small>"]
CO --> G1[":i-layers: <b>Group 1 — sliding window</b><br/><small>10 layers · last 32 tokens<br/>2 blocks</small>"]
CO --> G2[":i-layers: <b>Group 2 — sliding window</b><br/><small>10 layers · last 32 tokens<br/>2 blocks</small>"]
G0 --> POOL[(":i-database: <b>One BlockPool</b><br/><small>one page size for everyone</small>")]
G1 --> POOL
G2 --> POOL
class REQ neutral
class CO queue
class G0,G1,G2 compute
class POOL memoryHow it really works#
Layers that need less#
| Layer type | What it stores | Blocks needed for n tokens | Example models |
|---|---|---|---|
| Full attention | Keys and values for every token | ceil(n / block) | Llama, Qwen, Mistral |
| Sliding-window attention | Keys and values for the last w tokens only | about ceil(w / block), constant | Gemma 2 and 3, gpt-oss, Ministral |
| Chunked local attention | The current chunk only | Up to one chunk’s worth | Llama 4 |
| State-space (Mamba) | A fixed-size state, not per-token data | One, plus checkpoints | Jamba, Bamba |
| Latent attention (MLA) | One compressed vector per token instead of separate K and V per head | ceil(n / block), smaller pages | DeepSeek V3-class |
| Cross-attention | Keys and values of the encoder’s output | Fixed by the encoder length | Whisper |
Each attention layer in a model reports a spec saying which of these it is
(FullAttentionSpec, SlidingWindowSpec, MambaSpec, MLAAttentionSpec, … in
vllm/v1/kv_cache_interface.py). At startup the engine core collects the specs from the
workers and decides how to group them.
Why it matters: the sliding-window case#
Gemma-3-27B has 62 layers: 10 full attention and 52 with a sliding window of 1,024 tokens. For a 32,000-token prompt:
treat all 62 layers as full attention 62 × 32,000 = 1,984,000 layer-tokens
honour the window 10 × 32,000 + 52 × 1,024 = 373,248 layer-tokens
5.3 times less memoryThat is the difference between fitting three such requests and fitting sixteen.
The constraint: one page size#
The block pool hands out identical blocks. A block is a fixed number of bytes, a page. But a group of 10 layers and a group of 52 layers would need pages of different sizes for the same number of tokens. The design question is how to make every group use the same page size.
The answer is to make groups the same size. From the design document:
- A regular ratio. With 20 sliding-window and 10 full layers, form three groups of 10. Each
group’s page is
10 × block_size × bytes per token per layer. - No regular ratio. Gemma-3-27B’s 52 and 10 do not divide evenly. vLLM uses the smaller count as the group size: one full group of 10, five sliding-window groups of 10, and a seventh group with the 2 remaining layers and 8 layers of padding. The document is frank about it: “The solution has some memory waste and is not perfect.”
- Different sizes per token (state-space layers, whose state is much larger than one
token’s keys and values). The attention layers’
block_sizeis increased until one attention page is at least as large as the state, and the state is padded up to match. This is why some hybrid models run with blocks of several hundred or a thousand tokens rather than 16, and why--prefix-match-unitexists.
A request then has one block table per group. You can see it in the type of
NewRequestData.block_ids: tuple[list[int], ...], one list per group.
Layers share buffers across groups#
Physical memory is not one tensor per group. For n groups of m layers each, vLLM allocates
m buffers, and each buffer is shared by n layers, one from each group. Layer full.0,
layer sw.0 and layer sw.10 all read and write the same tensor, at different block indices.
Since the three groups hold disjoint sets of block IDs, they never collide. The benefit is that
“block 37” means the same thing for every layer: a slice of the same size at the same offset.
Freeing blocks as the window moves#
A sliding-window layer does not use a ring buffer. A ring would overwrite old tokens in place, and a block whose contents change cannot be shared or cached. Instead vLLM gives each token its own slot, exactly as for full attention, and frees blocks once they fall behind the window:
# Free the blocks that are skipped during the attention computation
# (e.g., tokens outside the sliding window).
self.coordinator.remove_skipped_blocks(request.request_id, ...)The freed positions in the request’s table are replaced by the null block (block 0), so the table keeps its shape and indices still line up with token positions. The freed blocks return to the pool with their hashes intact and may still serve another request’s cache hit.
Prefix caching across groups#
A cache hit is only usable if every layer can resume from it. The rules differ by type:
| Type | A prefix of length L is a valid hit if… |
|---|---|
| Full attention | every block from 0 to L is cached |
| Sliding window | the blocks covering the last w − 1 tokens before L are cached |
| State-space | a saved state exists exactly at L |
Full attention scans left to right and stops at the first miss. A sliding-window group can have a hit at a position where its earliest blocks are long gone, so it is scanned the other way, right to left, taking the first position that works.
The coordinator combines them:
- Find the longest hit for the full-attention group, scanning from the left.
- Within that length, find the longest hit for the other group, scanning from the right.
The result is valid for both. The design document notes the efficiency: the second scan starts from the first scan’s answer, so when there is no hit at all it stops immediately.
The cache key includes the group: internally the map is from (block hash, group id) to a
block. The same tokens are cached, and evicted, independently for each group.
Which coordinator you get#
get_kv_cache_coordinator chooses one at startup:
| Coordinator | When |
|---|---|
KVCacheCoordinatorNoPrefixCache | Prefix caching is disabled. Any number of groups; no lookups. |
UnitaryKVCacheCoordinator | Exactly one group. The common case; no intersection needed. |
HybridKVCacheCoordinator | Several groups with prefix caching on |
note
The published design document, written against an older commit, says the hybrid coordinator
“handles exactly two KV cache groups”. The class on main has grown since — it now deals
with state-space checkpoints, bounded replay for sliding windows and several specialised
layouts. Treat the document as the explanation of the idea and the source as the truth about
current limits.
Turning it off#
--disable-hybrid-kv-cache-manager makes vLLM treat every attention layer as full attention
for allocation, while the model still computes with its real window. It wastes memory —
this is the first row of the arithmetic above — but it removes a whole class of complexity,
and a few features have required it at one time or another. If the startup log says the hybrid
manager was disabled, your capacity is the full-attention figure.
KV sharing#
In some models a layer does not have its own cache at all; it reads another layer’s. vLLM
allocates nothing for such layers and points them at the owner’s buffer. They add nothing to
the per-token cost, which is why counting layers in config.json can overestimate the cache
for these architectures.
What this changes for you#
For Llama-style models: nothing. One group, one table, the arithmetic of Sizing the Cache.
For hybrid models:
- Capacity is not
blocks × 16. A request occupies blocks in several groups, a different number in each. vLLM computes a “group-aware” capacity and that is what the startup log prints asGPU KV cache size: N tokens. Trust that line over hand arithmetic. - Per-token cost is not constant. Short requests are relatively expensive (the window is not yet full); long ones are cheap per token.
- Block size may be large, so prefix-cache hits are coarse unless
--prefix-match-unitis set. - New architectures arrive with new specs. When a freshly released model is slow or memory-hungry in vLLM, the usual cause is that its layers are being treated as full attention until a dedicated spec lands.
Code#
Group formation and per-request memory for hybrid models, including the padding waste.
package main
import "fmt"
const blockSize = 16
func ceilDiv(a, b int) int { return (a + b - 1) / b }
// blocksFull: every token keeps its slot.
func blocksFull(n int) int { return ceilDiv(n, blockSize) }
// blocksWindow: slots for the last w tokens; whole blocks behind the window are freed.
func blocksWindow(n, w int) int {
skipped := n - w
if skipped < 0 {
skipped = 0
}
return ceilDiv(n, blockSize) - skipped/blockSize
}
type model struct {
name string
fullLayers int
windowLayers int
window int
groupSize int // layers per group = the smaller layer count
windowGroups int
paddingLayers int
}
func newModel(name string, full, win, window int) model {
m := model{name: name, fullLayers: full, windowLayers: win, window: window}
m.groupSize = min(full, win)
m.windowGroups = ceilDiv(win, m.groupSize)
m.paddingLayers = m.windowGroups*m.groupSize - win
return m
}
// cost returns layer-tokens of memory for one request of n tokens:
// blocks x block size x layers per group, summed over groups.
func (m model) cost(n int) (hybrid, allFull int) {
page := blockSize * m.groupSize
hybrid = blocksFull(n)*page*(m.fullLayers/m.groupSize) + blocksWindow(n, m.window)*page*m.windowGroups
allFull = blocksFull(n) * blockSize * (m.fullLayers + m.windowLayers)
return
}
func main() {
// The example from vLLM's design document: 10 full + 20 window layers, window 32.
toy := newModel("toy", 10, 20, 32)
fmt.Printf("%s: %d groups of %d layers (1 full, %d window), padding %d layers\n",
toy.name, 1+toy.windowGroups, toy.groupSize, toy.windowGroups, toy.paddingLayers)
fmt.Printf(" 112-token request: %d blocks for the full group, %d for each window group, %d in total\n\n",
blocksFull(112), blocksWindow(112, 32), blocksFull(112)+toy.windowGroups*blocksWindow(112, 32))
g := newModel("Gemma-3-27B", 10, 52, 1024)
fmt.Printf("%s: %d full + %d window layers -> 1 full group + %d window groups of %d, padding %d layers\n\n",
g.name, g.fullLayers, g.windowLayers, g.windowGroups, g.groupSize, g.paddingLayers)
fmt.Printf("%10s %16s %16s %10s\n", "tokens", "hybrid manager", "all-full", "saving")
for _, n := range []int{512, 2000, 8000, 32000, 128000} {
h, f := g.cost(n)
fmt.Printf("%10d %16d %16d %9.1fx\n", n, h, f, float64(f)/float64(h))
}
fmt.Println("\n(units: layer-tokens; multiply by bytes per token per layer for bytes)")
}Two things to read from the output. For a 512-token request the hybrid manager is worse than treating every layer as full attention: the window is not full yet, so nothing is saved, and the eight padding layers cost extra. The saving appears once requests exceed the window and grows with length. A model like this is a good fit for long-context work and a slightly poor one for short chats — which is exactly what its architecture was designed for.
Remember this#
- Layers that need different amounts of cache are placed in different KV cache groups.
- All groups draw equal-sized pages from one pool, so groups are made the same size, with padding if needed.
- A request has one block table per group.
- Sliding-window layers free blocks behind the window and leave null blocks in their place.
- A prefix-cache hit must be valid for every group: left-to-right for full attention, right-to-left for the rest.
- For hybrid models, read capacity from the startup log;
blocks × 16is wrong. --disable-hybrid-kv-cache-managertrades memory for simplicity.
Try it#
- Add a model with 30 full and 30 window layers and a window of 4,096. How many groups, how much padding, and what saving at 32,000 tokens?
- For Gemma-3-27B, find the request length at which the hybrid manager breaks even with all-full. Why is it a little above the window size and not exactly at it?
- Change
groupSizeto the greatest common divisor of the two layer counts instead of the minimum. For 52 and 10 that is 2. How many groups result, and what did vLLM gain by accepting padding instead?
Check yourself#
- Why must all KV cache groups have the same page size?
- Why does a sliding-window layer free old blocks instead of overwriting them in a ring?
- For a model with full and sliding-window layers, in which direction is each group scanned for a cache hit, and why?
Sources#
Checked on 5 October 2026 against main at commit 0c16eee.
- Hybrid KV cache manager (design) — the grouping cases and the Gemma-3-27B example
vllm/v1/kv_cache_interface.py— the spec classesvllm/v1/core/kv_cache_coordinator.pyvllm/v1/core/single_type_kv_cache_manager.py— one manager class per layer typevllm/v1/core/kv_cache_utils.py— group formation and page-size unification