The idea in one minute#
A decode step for a small model is thousands of tiny GPU operations, and Python needs a few
microseconds to launch each one. For a large model that overhead is lost in the arithmetic; for
a small one it can exceed the arithmetic. vLLM removes it in two ways. torch.compile turns
the model’s Python forward function into optimised code once, at startup. CUDA graphs then
record the entire sequence of GPU operations for one batch shape and replay it with a single
call. Both are done ahead of time for every batch size the scheduler might produce, which is
why a vLLM server takes a minute to start and never stutters afterwards. The price is startup
time, GPU memory, and a small amount of padding.
A picture#
flowchart TB
subgraph START["At startup, once"]
direction LR
C1[":python: <b>torch.compile</b><br/><small>trace the forward function,<br/>fuse ops, generate kernels</small>"] --> C2[":i-hard-drive: <b>cache on disk</b><br/><small>~/.cache/vllm/torch_compile_cache</small>"]
C2 --> C3[":nvidia: <b>capture CUDA graphs</b><br/><small>one per batch size:<br/>1, 2, 4, 8, 16, ... 512</small>"]
end
subgraph RUN["Every step"]
direction LR
B["batch of 37 tokens"] --> D{"dispatcher"}
D -->|"pure decode"| F["<b>FULL</b> graph, size 40<br/><small>whole forward pass, one launch</small>"]
D -->|"mixed or prefill"| P["<b>PIECEWISE</b> graphs, size 40<br/><small>everything except attention</small>"]
D -->|"no graph fits"| E["<b>eager</b><br/><small>op by op from Python</small>"]
end
START --> RUN
class C1,C3 compute
class C2 memory
class B neutral
class D queue
class F,P compute
class E warnHow it really works#
The overhead being removed#
Every PyTorch operation called from Python goes through the interpreter, argument checking, dispatch and a kernel launch. Call that 5 to 20 microseconds. A transformer layer is perhaps forty such operations; a 32-layer model has well over a thousand per forward pass:
1,300 operations × 10 µs ≈ 13 ms of pure overhead per step
decode step, large model, large batch arithmetic 30 ms overhead 13 ms noticeable
decode step, 1B model, batch of 4 arithmetic 3 ms overhead 13 ms dominantA CUDA graph is a recording of the GPU work submitted during one execution of a piece of code. Replaying it submits the same work with one call and no Python in between. The recording is rigid: same operations, same tensor shapes, same memory addresses. Only the contents of the input tensors may differ.
That rigidity fits vLLM well. Thanks to the flat batch (Building the Batch), the shape of a step is described by a single number, the total token count.
Compilation: torch.compile#
Before anything is recorded, vLLM compiles the model’s forward function. This is on by default: the design document calls it “a critical part of the framework”.
Three properties of vLLM’s use of it differ from ordinary torch.compile usage:
Everything is compiled before the first request. Normally torch.compile compiles lazily
and recompiles when it meets a new shape, which in a server would mean a multi-second stall on
some unlucky request. vLLM guarantees otherwise:
A unique aspect of vLLM’s
torch.compileintegration, is that we guarantee all the compilation finishes before we serve any requests. No requests will trigger new compilations.
It achieves this by compiling with a dynamic batch dimension and deliberately dropping the shape guards that would otherwise trigger recompilation.
The result is cached on disk. The first start of a given configuration compiles; later starts load the artefacts:
Using cache directory: ~/.cache/vllm/torch_compile_cache/1517964802/rank_0_0 for vLLM's torch.compileThe directory name is a hash of everything that could change the generated code: the model
configuration, the relevant vLLM configuration, the PyTorch configuration and the source of the
forward function itself. This is the purpose of the compute_hash() methods you see on every
config class, and of their warning comment: “Whenever a new field is added to this config,
ensure that it is included in the factors list if it affects the computation graph.” For
SchedulerConfig the hash covers max_num_batched_tokens and max_num_seqs, because buffer
sizes derived from them are baked into the compiled code.
tip
In a container, mount a volume at /root/.cache/vllm or bake the cache into the image.
Every pod that starts with an empty cache pays the full compilation time. The documentation
states plainly that you can “directly copy the whole ~/.cache/vllm/torch_compile_cache
directory in your deployment scenario to save a great amount of compilation time”.
The graph is split at attention. Attention kernels often cannot be recorded in a CUDA graph (their behaviour depends on the batch’s contents). vLLM’s compilation backend therefore splits the model’s graph at each attention call into pieces: the code between one attention and the next. This is piecewise compilation, and it is what makes the next section possible.
Compilation also applies vLLM’s own graph rewrites, called fusions: for example merging a normalisation with the quantisation that follows it into one kernel, or merging the all-reduce of tensor parallelism with the next normalisation.
CUDA graph modes#
CompilationConfig.cudagraph_mode (-cc.cudagraph_mode=…) has five values:
| Mode | What is recorded | When to use |
|---|---|---|
NONE | Nothing; eager execution | Debugging |
PIECEWISE | Each piece between attention calls. Attention runs eagerly. | Maximum compatibility; any attention backend |
FULL | The whole forward pass, attention included | Small models or short prompts, when the backend allows it |
FULL_DECODE_ONLY | The whole pass, but only for pure decode batches. Prefill and mixed batches run eagerly. | Decode-only instances in a disaggregated setup; saves the memory of the piecewise graphs |
FULL_AND_PIECEWISE | Both: full graphs for pure decode, piecewise for everything else | The default. “Generally the most performant setting … but also requires the most memory and takes the longest to capture.” |
A pure decode batch — every request contributing exactly one token, or exactly
1 + num_speculative_tokens — is what the code calls uniform. It is the most common step
by far on a busy server and the one where launch overhead is the largest share, which is why it
gets its own, fully recorded path.
If the attention backend cannot be recorded for a mode, vLLM downgrades silently to the closest
mode it can support. A backend that supports recording only for uniform batches turns FULL
into FULL_AND_PIECEWISE.
The dispatcher#
Each step, the CudagraphDispatcher receives a description of the batch:
class BatchDescriptor(NamedTuple):
num_tokens: int
num_reqs: int
uniform: bool = False
has_lora: bool = Falseand returns which recorded graph to use, or none. has_lora is part of the key because a batch
with adapters active executes different operations and needs its own recordings.
Which sizes are captured#
A graph is valid for exactly one token count, so vLLM records a ladder of sizes and pads each
batch up to the next rung. The default ladder, from VllmConfig._set_cudagraph_sizes:
cudagraph_capture_sizes = [i for i in [1, 2, 4] if i <= max_cudagraph_capture_size]
if max_cudagraph_capture_size >= 8:
# Step size 8 for small batch sizes, up to 256(not included)
cudagraph_capture_sizes += list(range(8, min(max_cudagraph_capture_size + 1, 256), 8))
if max_cudagraph_capture_size >= 256:
# Step size 16 for larger batch sizes
cudagraph_capture_sizes += list(range(256, max_cudagraph_capture_size + 1, 16))The top of the ladder defaults to:
max_cudagraph_capture_size = min(max_num_seqs × decode_query_len × 2, 512)
(1,024 on Blackwell data-centre GPUs)and never more than max_num_batched_tokens. Batches larger than the top rung — big prefill
chunks — run without a graph. That is fine: for a step of several thousand tokens the launch
overhead is a negligible fraction.
With max_num_seqs = 256 the ladder is 1, 2, 4, 8, 16, 24, …, 248, 256, 272, …, 512: 51
sizes. In the default mode each is captured twice (full and piecewise).
--performance-mode interactivity replaces the bottom of the ladder with every size from 1
to 32, so that a server with a handful of concurrent users never pads at all.
What it costs#
| Cost | Size | Mitigation |
|---|---|---|
| Startup time | Seconds to minutes: compilation on first start, graph capture on every start | Persist the compile cache; lower -O level; fewer capture sizes |
| GPU memory | Each recorded graph keeps its intermediate buffers. Hundreds of megabytes to a few gigabytes in total. | Since v0.21.0 this is estimated and subtracted from the KV cache budget (Sizing the Cache); reduce sizes or use FULL_DECODE_ONLY |
| Padding | A batch between two rungs computes a few extra rows | Negligible above ~16; use interactivity mode for tiny batches |
The startup log reports the second and first directly: Graph capturing finished in 9 secs, took 0.31 GiB.
Optimisation levels#
Four presets bundle these choices, selected like a compiler flag:
| Level | Compilation | CUDA graphs | Fusions | Use |
|---|---|---|---|---|
-O0 | None | NONE | None | Fastest startup; development and debugging |
-O1 | On | PIECEWISE | Basic | Fast startup with the main speedups |
-O2 | On, more compile ranges | FULL_AND_PIECEWISE | More | Default. Production. |
-O3 | Currently identical to -O2 | Reserved for slower or experimental optimisations |
vllm serve <model> -O1Any flag you set explicitly overrides the preset. -O0 is the first thing to try when a model
fails during startup: if it then works, the problem is in compilation or capture, not in the
model.
When it goes wrong#
| Symptom | Likely cause | First step |
|---|---|---|
| Startup takes many minutes every time | Compile cache not persisted between starts | Mount the cache directory |
| Out of memory during startup, after the KV cache is created | Graph capture needs more than was estimated | Lower --gpu-memory-utilization, or reduce capture sizes |
| Wrong output only with graphs on | A backend or custom operation is not capture-safe | Reproduce with -O0, then with -cc.cudagraph_mode=PIECEWISE |
| Fast at high concurrency, slow at 1 to 3 users | Padding and small-batch overhead | --performance-mode interactivity |
| A stale cache after patching model code | Hash did not cover the change | VLLM_DISABLE_COMPILE_CACHE=1 once |
Code#
The capture ladder, and the padding it causes.
package main
import "fmt"
// captureSizes reproduces the default ladder in VllmConfig._set_cudagraph_sizes.
func captureSizes(maxSize int, interactivity bool) []int {
var sizes []int
if interactivity {
for i := 1; i <= min(maxSize, 32); i++ {
sizes = append(sizes, i)
}
} else {
for _, i := range []int{1, 2, 4} {
if i <= maxSize {
sizes = append(sizes, i)
}
}
}
for i := 8; i < min(maxSize+1, 256); i += 8 {
if !contains(sizes, i) {
sizes = append(sizes, i)
}
}
for i := 256; i <= maxSize; i += 16 {
sizes = append(sizes, i)
}
return sizes
}
func contains(xs []int, x int) bool {
for _, v := range xs {
if v == x {
return true
}
}
return false
}
// padTo returns the smallest captured size >= n, or 0 if none (eager execution).
func padTo(sizes []int, n int) int {
best := 0
for _, s := range sizes {
if s >= n && (best == 0 || s < best) {
best = s
}
}
return best
}
func main() {
const maxNumSeqs = 256
maxSize := min(maxNumSeqs*1*2, 512) // decode_query_len = 1 without speculative decoding
def := captureSizes(maxSize, false)
inter := captureSizes(maxSize, true)
fmt.Printf("max_num_seqs=%d -> max capture size %d\n", maxNumSeqs, maxSize)
fmt.Printf("default ladder: %d sizes: %v ... %v\n", len(def), def[:8], def[len(def)-3:])
fmt.Printf("interactivity ladder: %d sizes\n\n", len(inter))
fmt.Printf("%8s %10s %8s %14s %8s\n", "batch", "default", "waste", "interactivity", "waste")
for _, n := range []int{1, 3, 5, 9, 17, 33, 100, 250, 300, 600} {
d, i := padTo(def, n), padTo(inter, n)
row := func(p int) (string, string) {
if p == 0 {
return "eager", "-"
}
return fmt.Sprint(p), fmt.Sprintf("%.0f%%", 100*float64(p-n)/float64(p))
}
ds, dw := row(d)
is, iw := row(i)
fmt.Printf("%8d %10s %8s %14s %8s\n", n, ds, dw, is, iw)
}
// Average waste if every batch size from 1 to maxSize were equally likely.
total := func(sizes []int) float64 {
w := 0.0
for n := 1; n <= maxSize; n++ {
p := padTo(sizes, n)
w += float64(p-n) / float64(p)
}
return 100 * w / float64(maxSize)
}
fmt.Printf("\naverage padding over all batch sizes: default %.1f%%, interactivity %.1f%%\n",
total(def), total(inter))
}The waste is largest for small batches — nine requests padded to sixteen is 44% — and is below 5% by a hundred. That is the whole case for the interactivity ladder: it spends capture time and memory on twenty-five extra sizes to remove padding exactly where it is proportionally worst.
Remember this#
- Launch overhead matters most for small models and small batches; CUDA graphs remove it by replaying a recording.
- A recording is tied to one token count, so vLLM captures a ladder of sizes and pads up.
torch.compileis on by default, finishes before any request is served, and is cached on disk.- The default mode is
FULL_AND_PIECEWISE: whole-pass graphs for pure decode, piecewise graphs otherwise. - The attention backend limits which modes are possible; vLLM downgrades automatically.
-O0to-O3trade startup time for speed;-O2is the default.- Persist
~/.cache/vllmor pay the compilation time on every cold start.
Try it#
- In the program, set
maxNumSeqsto 64. What is the top of the ladder, and what happens to a batch of 200 tokens? - Start the same model with
-O0,-O1and-O2. Record startup time and tokens per second at 4 concurrent requests. Which level gives the best value for a development loop? - Delete
~/.cache/vllm/torch_compile_cache, start a server, stop it, and start it again. Compare the twoinit engine … tooklines.
Check yourself#
- Why can a CUDA graph be reused for different requests but not for a different batch size?
- What makes a batch “uniform”, and why does that case get a fully recorded graph?
- Why does vLLM compile everything at startup instead of when a new shape first appears?
Sources#
Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.