The idea in one minute#
The number of blocks in the pool is fixed at startup and never changes. It is the result of one
subtraction and one division: take the share of the GPU you allowed vLLM to use, subtract what
the model’s weights and working memory need, and divide what remains by the size of one block.
That quotient decides how many tokens can be in flight at once, which decides how many requests
can run together, which decides whether your server preempts. You can compute it on paper
before buying a GPU, and you should, because the three numbers that drive it — layers, KV heads
and head size — are in the model’s config.json.
A picture#
flowchart LR
GPU[":nvidia: <b>GPU total memory</b><br/><small>e.g. 24 GiB</small>"] -->|"× gpu_memory_utilization (0.92)"| REQ["<b>requested</b><br/><small>22.1 GiB</small>"]
REQ --> SUB{"subtract"}
W[":i-hard-drive: weights"] --> SUB
A[":i-activity: peak activations<br/><small>one max-size dummy batch</small>"] --> SUB
NT[":i-package: non-torch<br/><small>CUDA context, NCCL</small>"] --> SUB
CG[":i-circuit-board: CUDA graph estimate"] --> SUB
SUB --> AV["<b>available for KV cache</b>"]
AV -->|"÷ bytes per block"| NB[":i-layers: <b>num_gpu_blocks</b>"]
NB -->|"× 16"| TOK["<b>capacity in tokens</b>"]
TOK -->|"÷ max_model_len"| CON["<b>maximum concurrency</b><br/><small>worst case</small>"]
class GPU,W,NB memory
class REQ,AV,TOK,CON neutral
class SUB queue
class A,NT,CG computeHow it really works#
Step 1: the budget#
When a worker starts it takes a snapshot of the GPU and computes its budget
(vllm/v1/worker/utils.py):
requested_memory = math.ceil(
init_snapshot.total_memory * cache_config.gpu_memory_utilization
)
if init_snapshot.free_memory < requested_memory:
raise ValueError(
f"Free memory on device ... on startup is less than desired GPU memory "
f"utilization ({cache_config.gpu_memory_utilization}, ...). Decrease GPU memory "
f"utilization or reduce GPU memory used by other processes."
)--gpu-memory-utilization defaults to 0.92. It is a fraction of the GPU’s total
memory, and it is a limit for this vLLM instance: weights, working memory and KV cache
together. It does not know about other processes. Two vLLM servers sharing one GPU must each be
given their own share, for example 0.45 each.
If another process already occupies more than 8% of the card, the default fails at startup with the error above. The fix is a lower utilisation, not a retry.
Step 2: measure what the model needs#
The worker loads the weights, then runs one forward pass with dummy inputs of the largest
size the scheduler will ever send, and watches memory (vllm/utils/mem_utils.py). The
profiler’s docstring classifies GPU memory into three kinds and works through an example:
1. memory used by anything other than the current vLLM instance
2. memory used by torch in the current vLLM instance
3. memory used in the current vLLM instance, but not by torch
non-KV-cache memory =
a. model weights (kind 2)
+ b. peak activation tensors (kind 2, the increase in torch's peak during the dummy pass)
+ c. non-torch components (kind 3: CUDA context, NCCL, attention-backend buffers)The dummy pass matters. Activation memory grows with the batch, and the largest batch is
max_num_batched_tokens tokens. Raising that flag raises (b) and so shrinks the KV cache.
This is one of the less obvious trade-offs in vLLM: a larger token budget costs memory twice,
once directly and once in blocks you no longer have.
A fourth term was added later. CUDA graphs — recorded execution plans that make small batches fast — also consume memory, and they are captured after the KV cache is sized. vLLM now estimates their cost first and subtracts it:
self.available_kv_cache_memory_bytes = (
self.requested_memory
- profile_result.non_kv_cache_memory
- cudagraph_memory_estimate_applied
)The startup log is explicit that this changed behaviour. With example numbers filled in, it
reads: “CUDA graph memory profiling is enabled (default since v0.21.0). The current
--gpu-memory-utilization=0.9200 is equivalent to --gpu-memory-utilization=0.9050 without
CUDA graph memory profiling.” It also tells you which value to use if you want the cache size
you had before.
The result appears as Available KV cache memory: 18.4 GiB.
Step 3: the size of one block#
Each attention layer reports what its cache needs as a KV cache spec. For ordinary
attention (AttentionSpec in vllm/v1/kv_cache_interface.py):
@property
def state_content_size_bytes(self) -> int:
return (self.head_size + self.head_size_v) * get_dtype_size(self.dtype)
@property
def unpadded_page_size_bytes(self) -> int:
return self.num_heads * self.num_states * self.state_content_size_bytesIn plain terms, for one layer and one block:
page size = KV heads × tokens per block × (key size + value size) × bytes per number
8 × 16 × (128 + 128) × 2
= 65,536 bytes for one layer of Llama-3.1-8Band a block spans every layer:
bytes per block = layers × page size = 32 × 65,536 = 2,097,152 bytes = 2 MiB
bytes per token = bytes per block / 16 = 131,072 bytes = 128 KiBThe inputs come from the model’s config.json:
| Quantity | config.json field | Llama-3.1-8B |
|---|---|---|
| Layers | num_hidden_layers | 32 |
| KV heads | num_key_value_heads | 8 |
| Head size | hidden_size / num_attention_heads | 4096 / 32 = 128 |
| Bytes per number | torch_dtype (bfloat16 → 2) | 2 |
important
Use num_key_value_heads, not num_attention_heads. Most current models use grouped-query
attention, where many query heads share few key/value heads. Llama-3.1-8B has 32 attention
heads but only 8 KV heads, so its cache is a quarter of what the wrong field suggests.
Step 4: divide#
num_blocks = available_memory // bytes_per_block
num_blocks = may_override_num_blocks(vllm_config, num_blocks)One contiguous tensor per layer is then allocated on the GPU with room for num_blocks pages.
The number is sent to the engine core, which builds the BlockPool from
Blocks and the Pool around it.
--num-gpu-blocks-override N replaces the computed value. The config describes its purpose in
four words: “Used for testing preemption.” Do not use it to get more blocks than fit; the
memory still has to exist.
Step 5: the sanity check, and the log line#
Before allocating, vLLM checks that at least one request of maximum length fits:
ValueError: To serve at least one request with the model's max seq len (131072),
(16.00 GiB KV cache is needed, which is larger than the available KV cache memory
(6.31 GiB). Based on the available memory, the estimated maximum model length is 51440.
Try increasing `gpu_memory_utilization` ... or decreasing `max_model_len` ...This is the most common startup failure for a model with a long context window on a modest GPU.
The message contains its own solution: pass --max-model-len at or below the estimate.
Or let vLLM choose. With --max-model-len -1, a binary search finds the largest context that
fits and uses it, logging “Auto-fit max_model_len: reduced from 131072 to 51440 to fit in
available GPU memory”. The reduced value is sent back to the API server in the engine’s ready
message, so request validation uses the real limit.
On success:
GPU KV cache size: 688,304 tokens, Maximum concurrency for 32,768 tokens per request: 21.01xRead those two numbers differently.
- Tokens is the hard capacity: the sum of prompt and generated tokens over all running requests cannot exceed it. (Shared prefix blocks count once.)
- Maximum concurrency assumes every request uses the whole context window. Real requests are far shorter. If yours average 2,000 tokens, this cache holds about 340 of them, not 21.
A useful planning rule: capacity ÷ p95 total tokens per request should comfortably exceed
max_num_seqs. If it does not, you will see preemption
(Priorities, Preemption and Queues).
Tensor parallelism changes the arithmetic#
With --tensor-parallel-size N, attention heads are split across the GPUs. Each GPU stores the
KV for its own share of the KV heads, so:
bytes per token, per GPU = (single-GPU value) / Nand each GPU has its own memory budget. Splitting a model over more GPUs therefore increases cache capacity twice over: the weights per GPU shrink, leaving more room, and each token costs less on each card. The calculator below shows a 70B model going from barely usable on two 80 GiB GPUs to comfortable on four.
The knobs, ranked by effect#
| Knob | Effect on capacity | Cost |
|---|---|---|
| Quantised weights (AWQ, GPTQ, FP8…) | Large: frees weight memory for cache | Some quality; see Quantization |
--kv-cache-dtype fp8 | ×2: one byte per number instead of two | Small quality loss; backend support varies |
| More GPUs with tensor parallelism | Large | Hardware; communication overhead |
--gpu-memory-utilization 0.95 | A few percent | Less safety margin against out-of-memory |
Lower --max-num-batched-tokens | Shrinks peak activations | Slower prefill |
Lower --max-num-seqs | Shrinks per-request buffers and CUDA graphs | Less concurrency |
--max-model-len | None on capacity in tokens; raises “maximum concurrency” and lets startup pass | Long requests rejected |
--kv-cache-memory-bytes N | Sets cache size directly, ignoring the utilisation fraction | You own the risk of running out |
--kv-cache-memory-bytes skips the subtraction entirely: the log says it “skipped memory
profiling. This does not respect the gpu_memory_utilization config.” It is the right tool
when several processes share a GPU and you want each to have an exact, predictable cache.
What does not count#
Two things people expect to be in this budget are not:
- Other processes on the GPU. vLLM measures them once at startup and assumes they do not change. If another process grows afterwards, vLLM can fail later with an out-of-memory error even though its own usage is constant. The profiler asserts this assumption and reports “This happens when other processes sharing the same container release GPU memory while vLLM is profiling during initialization.”
- Growth after startup. There is none by design. If memory use climbs while serving, something outside the KV cache is allocating: typically an attention backend’s workspace hitting a new maximum, or a LoRA adapter being loaded.
Code#
The whole calculation for three models. Change the numbers to match your hardware before you rent it.
package main
import "fmt"
const (
GiB = 1 << 30
blockSize = 16
)
type model struct {
name string
layers int
kvHeads int // num_key_value_heads, not num_attention_heads
headDim int // hidden_size / num_attention_heads
weightGiB float64 // as loaded
maxLen int
}
// bytesPerToken is the KV cache cost of one token on one GPU:
// layers x KV heads x (key + value) x bytes per element.
func (m model) bytesPerToken(dtypeBytes, tp int) int {
return m.layers * (m.kvHeads / tp) * (2 * m.headDim) * dtypeBytes
}
func size(m model, gpuGiB, util float64, dtypeBytes, tp int, otherGiB float64) {
requested := gpuGiB * util
available := requested - m.weightGiB/float64(tp) - otherGiB
perBlock := m.bytesPerToken(dtypeBytes, tp) * blockSize
blocks := int(available * GiB / float64(perBlock))
tokens := blocks * blockSize
fmt.Printf("%-14s %3.0f GiB x%d util %.2f kv %2d-bit | %5.1f KiB/token %5.1f GiB for KV %8d tokens %5.1fx at %d\n",
m.name, gpuGiB, tp, util, dtypeBytes*8,
float64(m.bytesPerToken(dtypeBytes, tp))/1024, available, tokens,
float64(tokens)/float64(m.maxLen), m.maxLen)
}
func main() {
qwen := model{"Qwen2.5-1.5B", 28, 2, 128, 2.9, 32768}
llama8 := model{"Llama-3.1-8B", 32, 8, 128, 15.0, 131072}
llama70 := model{"Llama-3.3-70B", 80, 8, 128, 131.5, 131072}
const other = 0.8 // peak activations + non-torch memory + CUDA graphs, GiB per GPU
fmt.Println("Small model, small GPU:")
size(qwen, 24, 0.92, 2, 1, other)
fmt.Println("\nSame model and GPU, two knobs:")
size(qwen, 24, 0.50, 2, 1, other)
size(qwen, 24, 0.92, 1, 1, other) // --kv-cache-dtype fp8
fmt.Println("\n8B model:")
size(llama8, 24, 0.92, 2, 1, other)
size(llama8, 80, 0.92, 2, 1, other)
fmt.Println("\n70B model: does not fit one 80 GiB GPU; tensor parallel over 2 and 4:")
size(llama70, 80, 0.92, 2, 2, other)
size(llama70, 80, 0.92, 2, 4, other)
}Two rows print a concurrency below 1.0. Those are configurations vLLM refuses to start: not
even one full-length request fits. Both become usable the moment --max-model-len is lowered —
the capacity in tokens is unchanged, but the check passes and the server can admit many shorter
requests.
The value of other is a guess; it is the one input you must measure. Start the server once
and read Available KV cache memory from the log, then solve for it.
Remember this#
num_blocks = (total × utilisation − weights − peak activations − non-torch − CUDA graph estimate) ÷ bytes per block.- Bytes per token = layers × KV heads × 2 × head size × bytes per number.
--gpu-memory-utilizationdefaults to 0.92 of total memory and covers everything this instance uses.- The pool is sized once; it never grows.
- Capacity in tokens is the real limit. “Maximum concurrency” is the worst case at full context length.
--max-model-lendoes not change capacity; it changes whether startup succeeds and how many requests the worst case allows.- FP8 KV cache doubles capacity; tensor parallelism raises it twice over.
- A larger
max_num_batched_tokensshrinks the cache by raising peak activation memory.
Try it#
- Find
num_hidden_layers,num_key_value_heads,hidden_sizeandnum_attention_headsin theconfig.jsonof a model you use. Add it to the program. How many tokens fit on your GPU? - For Llama-3.1-8B on 24 GiB, what
--max-model-lenmakes the worst-case concurrency exactly 8? Check it against the “estimated maximum model length” logic. - Start a real server twice, with
--max-num-batched-tokens 2048and16384, and compareAvailable KV cache memory. The difference is the extra peak activation memory.
Check yourself#
- Which
config.jsonfield gives the number of heads that matters for the KV cache, and why is it often smaller than the number of attention heads? - vLLM refuses to start, saying the KV cache needed is larger than what is available. Name two flags that fix it and what each gives up.
- Why does raising
max_num_batched_tokensreduce the number of KV blocks?
Sources#
Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.
vllm/v1/worker/gpu_worker.py—determine_available_memoryvllm/utils/mem_utils.py—memory_profilingvllm/v1/kv_cache_interface.py—AttentionSpec.page_size_bytesvllm/v1/core/kv_cache_utils.py— block count, the capacity log line, auto-fitvllm/config/cache.py- Conserving memory