Pidoku

Sizing the Cache: From Gigabytes to Blocks

Intermediate 50 min Difficulty 3/5 Lesson 04 of 05

Prerequisites Prefix Caching

The idea in one minute#

The number of blocks in the pool is fixed at startup and never changes. It is the result of one subtraction and one division: take the share of the GPU you allowed vLLM to use, subtract what the model’s weights and working memory need, and divide what remains by the size of one block. That quotient decides how many tokens can be in flight at once, which decides how many requests can run together, which decides whether your server preempts. You can compute it on paper before buying a GPU, and you should, because the three numbers that drive it — layers, KV heads and head size — are in the model’s config.json.

A picture#

flowchart LR
  GPU[":nvidia: <b>GPU total memory</b><br/><small>e.g. 24 GiB</small>"] -->|"× gpu_memory_utilization (0.92)"| REQ["<b>requested</b><br/><small>22.1 GiB</small>"]
  REQ --> SUB{"subtract"}
  W[":i-hard-drive: weights"] --> SUB
  A[":i-activity: peak activations<br/><small>one max-size dummy batch</small>"] --> SUB
  NT[":i-package: non-torch<br/><small>CUDA context, NCCL</small>"] --> SUB
  CG[":i-circuit-board: CUDA graph estimate"] --> SUB
  SUB --> AV["<b>available for KV cache</b>"]
  AV -->|"÷ bytes per block"| NB[":i-layers: <b>num_gpu_blocks</b>"]
  NB -->|"× 16"| TOK["<b>capacity in tokens</b>"]
  TOK -->|"÷ max_model_len"| CON["<b>maximum concurrency</b><br/><small>worst case</small>"]
  class GPU,W,NB memory
  class REQ,AV,TOK,CON neutral
  class SUB queue
  class A,NT,CG compute

How it really works#

Step 1: the budget#

When a worker starts it takes a snapshot of the GPU and computes its budget (vllm/v1/worker/utils.py):

Python
requested_memory = math.ceil(
    init_snapshot.total_memory * cache_config.gpu_memory_utilization
)
if init_snapshot.free_memory < requested_memory:
    raise ValueError(
        f"Free memory on device ... on startup is less than desired GPU memory "
        f"utilization ({cache_config.gpu_memory_utilization}, ...). Decrease GPU memory "
        f"utilization or reduce GPU memory used by other processes."
    )

--gpu-memory-utilization defaults to 0.92. It is a fraction of the GPU’s total memory, and it is a limit for this vLLM instance: weights, working memory and KV cache together. It does not know about other processes. Two vLLM servers sharing one GPU must each be given their own share, for example 0.45 each.

If another process already occupies more than 8% of the card, the default fails at startup with the error above. The fix is a lower utilisation, not a retry.

Step 2: measure what the model needs#

The worker loads the weights, then runs one forward pass with dummy inputs of the largest size the scheduler will ever send, and watches memory (vllm/utils/mem_utils.py). The profiler’s docstring classifies GPU memory into three kinds and works through an example:

1. memory used by anything other than the current vLLM instance
2. memory used by torch in the current vLLM instance
3. memory used in the current vLLM instance, but not by torch

non-KV-cache memory =
     a. model weights                         (kind 2)
   + b. peak activation tensors               (kind 2, the increase in torch's peak during the dummy pass)
   + c. non-torch components                  (kind 3: CUDA context, NCCL, attention-backend buffers)

The dummy pass matters. Activation memory grows with the batch, and the largest batch is max_num_batched_tokens tokens. Raising that flag raises (b) and so shrinks the KV cache. This is one of the less obvious trade-offs in vLLM: a larger token budget costs memory twice, once directly and once in blocks you no longer have.

A fourth term was added later. CUDA graphs — recorded execution plans that make small batches fast — also consume memory, and they are captured after the KV cache is sized. vLLM now estimates their cost first and subtracts it:

Python
self.available_kv_cache_memory_bytes = (
    self.requested_memory
    - profile_result.non_kv_cache_memory
    - cudagraph_memory_estimate_applied
)

The startup log is explicit that this changed behaviour. With example numbers filled in, it reads: “CUDA graph memory profiling is enabled (default since v0.21.0). The current --gpu-memory-utilization=0.9200 is equivalent to --gpu-memory-utilization=0.9050 without CUDA graph memory profiling.” It also tells you which value to use if you want the cache size you had before.

The result appears as Available KV cache memory: 18.4 GiB.

Step 3: the size of one block#

Each attention layer reports what its cache needs as a KV cache spec. For ordinary attention (AttentionSpec in vllm/v1/kv_cache_interface.py):

Python
@property
def state_content_size_bytes(self) -> int:
    return (self.head_size + self.head_size_v) * get_dtype_size(self.dtype)

@property
def unpadded_page_size_bytes(self) -> int:
    return self.num_heads * self.num_states * self.state_content_size_bytes

In plain terms, for one layer and one block:

page size = KV heads × tokens per block × (key size + value size) × bytes per number

            8      ×       16          ×    (128 + 128)          ×        2
          = 65,536 bytes for one layer of Llama-3.1-8B

and a block spans every layer:

bytes per block = layers × page size  = 32 × 65,536 = 2,097,152 bytes = 2 MiB
bytes per token = bytes per block / 16 = 131,072 bytes = 128 KiB

The inputs come from the model’s config.json:

Quantityconfig.json fieldLlama-3.1-8B
Layersnum_hidden_layers32
KV headsnum_key_value_heads8
Head sizehidden_size / num_attention_heads4096 / 32 = 128
Bytes per numbertorch_dtype (bfloat16 → 2)2

important

Use num_key_value_heads, not num_attention_heads. Most current models use grouped-query attention, where many query heads share few key/value heads. Llama-3.1-8B has 32 attention heads but only 8 KV heads, so its cache is a quarter of what the wrong field suggests.

Step 4: divide#

Python
num_blocks = available_memory // bytes_per_block
num_blocks = may_override_num_blocks(vllm_config, num_blocks)

One contiguous tensor per layer is then allocated on the GPU with room for num_blocks pages. The number is sent to the engine core, which builds the BlockPool from Blocks and the Pool around it.

--num-gpu-blocks-override N replaces the computed value. The config describes its purpose in four words: “Used for testing preemption.” Do not use it to get more blocks than fit; the memory still has to exist.

Step 5: the sanity check, and the log line#

Before allocating, vLLM checks that at least one request of maximum length fits:

ValueError: To serve at least one request with the model's max seq len (131072),
(16.00 GiB KV cache is needed, which is larger than the available KV cache memory
(6.31 GiB). Based on the available memory, the estimated maximum model length is 51440.
Try increasing `gpu_memory_utilization` ... or decreasing `max_model_len` ...

This is the most common startup failure for a model with a long context window on a modest GPU. The message contains its own solution: pass --max-model-len at or below the estimate.

Or let vLLM choose. With --max-model-len -1, a binary search finds the largest context that fits and uses it, logging “Auto-fit max_model_len: reduced from 131072 to 51440 to fit in available GPU memory”. The reduced value is sent back to the API server in the engine’s ready message, so request validation uses the real limit.

On success:

GPU KV cache size: 688,304 tokens, Maximum concurrency for 32,768 tokens per request: 21.01x

Read those two numbers differently.

  • Tokens is the hard capacity: the sum of prompt and generated tokens over all running requests cannot exceed it. (Shared prefix blocks count once.)
  • Maximum concurrency assumes every request uses the whole context window. Real requests are far shorter. If yours average 2,000 tokens, this cache holds about 340 of them, not 21.

A useful planning rule: capacity ÷ p95 total tokens per request should comfortably exceed max_num_seqs. If it does not, you will see preemption (Priorities, Preemption and Queues).

Tensor parallelism changes the arithmetic#

With --tensor-parallel-size N, attention heads are split across the GPUs. Each GPU stores the KV for its own share of the KV heads, so:

bytes per token, per GPU = (single-GPU value) / N

and each GPU has its own memory budget. Splitting a model over more GPUs therefore increases cache capacity twice over: the weights per GPU shrink, leaving more room, and each token costs less on each card. The calculator below shows a 70B model going from barely usable on two 80 GiB GPUs to comfortable on four.

The knobs, ranked by effect#

KnobEffect on capacityCost
Quantised weights (AWQ, GPTQ, FP8…)Large: frees weight memory for cacheSome quality; see Quantization
--kv-cache-dtype fp8×2: one byte per number instead of twoSmall quality loss; backend support varies
More GPUs with tensor parallelismLargeHardware; communication overhead
--gpu-memory-utilization 0.95A few percentLess safety margin against out-of-memory
Lower --max-num-batched-tokensShrinks peak activationsSlower prefill
Lower --max-num-seqsShrinks per-request buffers and CUDA graphsLess concurrency
--max-model-lenNone on capacity in tokens; raises “maximum concurrency” and lets startup passLong requests rejected
--kv-cache-memory-bytes NSets cache size directly, ignoring the utilisation fractionYou own the risk of running out

--kv-cache-memory-bytes skips the subtraction entirely: the log says it “skipped memory profiling. This does not respect the gpu_memory_utilization config.” It is the right tool when several processes share a GPU and you want each to have an exact, predictable cache.

What does not count#

Two things people expect to be in this budget are not:

  • Other processes on the GPU. vLLM measures them once at startup and assumes they do not change. If another process grows afterwards, vLLM can fail later with an out-of-memory error even though its own usage is constant. The profiler asserts this assumption and reports “This happens when other processes sharing the same container release GPU memory while vLLM is profiling during initialization.”
  • Growth after startup. There is none by design. If memory use climbs while serving, something outside the KV cache is allocating: typically an attention backend’s workspace hitting a new maximum, or a LoRA adapter being loaded.

Code#

The whole calculation for three models. Change the numbers to match your hardware before you rent it.

Go
package main

import "fmt"

const (
	GiB       = 1 << 30
	blockSize = 16
)

type model struct {
	name      string
	layers    int
	kvHeads   int     // num_key_value_heads, not num_attention_heads
	headDim   int     // hidden_size / num_attention_heads
	weightGiB float64 // as loaded
	maxLen    int
}

// bytesPerToken is the KV cache cost of one token on one GPU:
// layers x KV heads x (key + value) x bytes per element.
func (m model) bytesPerToken(dtypeBytes, tp int) int {
	return m.layers * (m.kvHeads / tp) * (2 * m.headDim) * dtypeBytes
}

func size(m model, gpuGiB, util float64, dtypeBytes, tp int, otherGiB float64) {
	requested := gpuGiB * util
	available := requested - m.weightGiB/float64(tp) - otherGiB
	perBlock := m.bytesPerToken(dtypeBytes, tp) * blockSize
	blocks := int(available * GiB / float64(perBlock))
	tokens := blocks * blockSize
	fmt.Printf("%-14s %3.0f GiB x%d  util %.2f  kv %2d-bit | %5.1f KiB/token  %5.1f GiB for KV  %8d tokens  %5.1fx at %d\n",
		m.name, gpuGiB, tp, util, dtypeBytes*8,
		float64(m.bytesPerToken(dtypeBytes, tp))/1024, available, tokens,
		float64(tokens)/float64(m.maxLen), m.maxLen)
}

func main() {
	qwen := model{"Qwen2.5-1.5B", 28, 2, 128, 2.9, 32768}
	llama8 := model{"Llama-3.1-8B", 32, 8, 128, 15.0, 131072}
	llama70 := model{"Llama-3.3-70B", 80, 8, 128, 131.5, 131072}

	const other = 0.8 // peak activations + non-torch memory + CUDA graphs, GiB per GPU

	fmt.Println("Small model, small GPU:")
	size(qwen, 24, 0.92, 2, 1, other)

	fmt.Println("\nSame model and GPU, two knobs:")
	size(qwen, 24, 0.50, 2, 1, other)
	size(qwen, 24, 0.92, 1, 1, other) // --kv-cache-dtype fp8

	fmt.Println("\n8B model:")
	size(llama8, 24, 0.92, 2, 1, other)
	size(llama8, 80, 0.92, 2, 1, other)

	fmt.Println("\n70B model: does not fit one 80 GiB GPU; tensor parallel over 2 and 4:")
	size(llama70, 80, 0.92, 2, 2, other)
	size(llama70, 80, 0.92, 2, 4, other)
}

Two rows print a concurrency below 1.0. Those are configurations vLLM refuses to start: not even one full-length request fits. Both become usable the moment --max-model-len is lowered — the capacity in tokens is unchanged, but the check passes and the server can admit many shorter requests.

The value of other is a guess; it is the one input you must measure. Start the server once and read Available KV cache memory from the log, then solve for it.

Remember this#

  • num_blocks = (total × utilisation − weights − peak activations − non-torch − CUDA graph estimate) ÷ bytes per block.
  • Bytes per token = layers × KV heads × 2 × head size × bytes per number.
  • --gpu-memory-utilization defaults to 0.92 of total memory and covers everything this instance uses.
  • The pool is sized once; it never grows.
  • Capacity in tokens is the real limit. “Maximum concurrency” is the worst case at full context length.
  • --max-model-len does not change capacity; it changes whether startup succeeds and how many requests the worst case allows.
  • FP8 KV cache doubles capacity; tensor parallelism raises it twice over.
  • A larger max_num_batched_tokens shrinks the cache by raising peak activation memory.

Try it#

  1. Find num_hidden_layers, num_key_value_heads, hidden_size and num_attention_heads in the config.json of a model you use. Add it to the program. How many tokens fit on your GPU?
  2. For Llama-3.1-8B on 24 GiB, what --max-model-len makes the worst-case concurrency exactly 8? Check it against the “estimated maximum model length” logic.
  3. Start a real server twice, with --max-num-batched-tokens 2048 and 16384, and compare Available KV cache memory. The difference is the extra peak activation memory.

Check yourself#

  1. Which config.json field gives the number of heads that matters for the KV cache, and why is it often smaller than the number of attention heads?
  2. vLLM refuses to start, saying the KV cache needed is larger than what is available. Name two flags that fix it and what each gives up.
  3. Why does raising max_num_batched_tokens reduce the number of KV blocks?

Sources#

Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom