Pidoku

CUDA Graphs and Compilation

Advanced 50 min Difficulty 4/5 Lesson 04 of 05

Prerequisites Attention Backends

The idea in one minute#

A decode step for a small model is thousands of tiny GPU operations, and Python needs a few microseconds to launch each one. For a large model that overhead is lost in the arithmetic; for a small one it can exceed the arithmetic. vLLM removes it in two ways. torch.compile turns the model’s Python forward function into optimised code once, at startup. CUDA graphs then record the entire sequence of GPU operations for one batch shape and replay it with a single call. Both are done ahead of time for every batch size the scheduler might produce, which is why a vLLM server takes a minute to start and never stutters afterwards. The price is startup time, GPU memory, and a small amount of padding.

A picture#

flowchart TB
  subgraph START["At startup, once"]
    direction LR
    C1[":python: <b>torch.compile</b><br/><small>trace the forward function,<br/>fuse ops, generate kernels</small>"] --> C2[":i-hard-drive: <b>cache on disk</b><br/><small>~/.cache/vllm/torch_compile_cache</small>"]
    C2 --> C3[":nvidia: <b>capture CUDA graphs</b><br/><small>one per batch size:<br/>1, 2, 4, 8, 16, ... 512</small>"]
  end
  subgraph RUN["Every step"]
    direction LR
    B["batch of 37 tokens"] --> D{"dispatcher"}
    D -->|"pure decode"| F["<b>FULL</b> graph, size 40<br/><small>whole forward pass, one launch</small>"]
    D -->|"mixed or prefill"| P["<b>PIECEWISE</b> graphs, size 40<br/><small>everything except attention</small>"]
    D -->|"no graph fits"| E["<b>eager</b><br/><small>op by op from Python</small>"]
  end
  START --> RUN
  class C1,C3 compute
  class C2 memory
  class B neutral
  class D queue
  class F,P compute
  class E warn

How it really works#

The overhead being removed#

Every PyTorch operation called from Python goes through the interpreter, argument checking, dispatch and a kernel launch. Call that 5 to 20 microseconds. A transformer layer is perhaps forty such operations; a 32-layer model has well over a thousand per forward pass:

1,300 operations × 10 µs  ≈  13 ms of pure overhead per step

decode step, large model, large batch     arithmetic 30 ms    overhead 13 ms   noticeable
decode step, 1B model, batch of 4         arithmetic  3 ms    overhead 13 ms   dominant

A CUDA graph is a recording of the GPU work submitted during one execution of a piece of code. Replaying it submits the same work with one call and no Python in between. The recording is rigid: same operations, same tensor shapes, same memory addresses. Only the contents of the input tensors may differ.

That rigidity fits vLLM well. Thanks to the flat batch (Building the Batch), the shape of a step is described by a single number, the total token count.

Compilation: torch.compile#

Before anything is recorded, vLLM compiles the model’s forward function. This is on by default: the design document calls it “a critical part of the framework”.

Three properties of vLLM’s use of it differ from ordinary torch.compile usage:

Everything is compiled before the first request. Normally torch.compile compiles lazily and recompiles when it meets a new shape, which in a server would mean a multi-second stall on some unlucky request. vLLM guarantees otherwise:

A unique aspect of vLLM’s torch.compile integration, is that we guarantee all the compilation finishes before we serve any requests. No requests will trigger new compilations.

It achieves this by compiling with a dynamic batch dimension and deliberately dropping the shape guards that would otherwise trigger recompilation.

The result is cached on disk. The first start of a given configuration compiles; later starts load the artefacts:

Using cache directory: ~/.cache/vllm/torch_compile_cache/1517964802/rank_0_0 for vLLM's torch.compile

The directory name is a hash of everything that could change the generated code: the model configuration, the relevant vLLM configuration, the PyTorch configuration and the source of the forward function itself. This is the purpose of the compute_hash() methods you see on every config class, and of their warning comment: “Whenever a new field is added to this config, ensure that it is included in the factors list if it affects the computation graph.” For SchedulerConfig the hash covers max_num_batched_tokens and max_num_seqs, because buffer sizes derived from them are baked into the compiled code.

tip

In a container, mount a volume at /root/.cache/vllm or bake the cache into the image. Every pod that starts with an empty cache pays the full compilation time. The documentation states plainly that you can “directly copy the whole ~/.cache/vllm/torch_compile_cache directory in your deployment scenario to save a great amount of compilation time”.

The graph is split at attention. Attention kernels often cannot be recorded in a CUDA graph (their behaviour depends on the batch’s contents). vLLM’s compilation backend therefore splits the model’s graph at each attention call into pieces: the code between one attention and the next. This is piecewise compilation, and it is what makes the next section possible.

Compilation also applies vLLM’s own graph rewrites, called fusions: for example merging a normalisation with the quantisation that follows it into one kernel, or merging the all-reduce of tensor parallelism with the next normalisation.

CUDA graph modes#

CompilationConfig.cudagraph_mode (-cc.cudagraph_mode=…) has five values:

ModeWhat is recordedWhen to use
NONENothing; eager executionDebugging
PIECEWISEEach piece between attention calls. Attention runs eagerly.Maximum compatibility; any attention backend
FULLThe whole forward pass, attention includedSmall models or short prompts, when the backend allows it
FULL_DECODE_ONLYThe whole pass, but only for pure decode batches. Prefill and mixed batches run eagerly.Decode-only instances in a disaggregated setup; saves the memory of the piecewise graphs
FULL_AND_PIECEWISEBoth: full graphs for pure decode, piecewise for everything elseThe default. “Generally the most performant setting … but also requires the most memory and takes the longest to capture.”

A pure decode batch — every request contributing exactly one token, or exactly 1 + num_speculative_tokens — is what the code calls uniform. It is the most common step by far on a busy server and the one where launch overhead is the largest share, which is why it gets its own, fully recorded path.

If the attention backend cannot be recorded for a mode, vLLM downgrades silently to the closest mode it can support. A backend that supports recording only for uniform batches turns FULL into FULL_AND_PIECEWISE.

The dispatcher#

Each step, the CudagraphDispatcher receives a description of the batch:

Python
class BatchDescriptor(NamedTuple):
    num_tokens: int
    num_reqs: int
    uniform: bool = False
    has_lora: bool = False

and returns which recorded graph to use, or none. has_lora is part of the key because a batch with adapters active executes different operations and needs its own recordings.

Which sizes are captured#

A graph is valid for exactly one token count, so vLLM records a ladder of sizes and pads each batch up to the next rung. The default ladder, from VllmConfig._set_cudagraph_sizes:

Python
cudagraph_capture_sizes = [i for i in [1, 2, 4] if i <= max_cudagraph_capture_size]
if max_cudagraph_capture_size >= 8:
    # Step size 8 for small batch sizes, up to 256(not included)
    cudagraph_capture_sizes += list(range(8, min(max_cudagraph_capture_size + 1, 256), 8))
if max_cudagraph_capture_size >= 256:
    # Step size 16 for larger batch sizes
    cudagraph_capture_sizes += list(range(256, max_cudagraph_capture_size + 1, 16))

The top of the ladder defaults to:

max_cudagraph_capture_size = min(max_num_seqs × decode_query_len × 2,  512)
                                                                       (1,024 on Blackwell data-centre GPUs)

and never more than max_num_batched_tokens. Batches larger than the top rung — big prefill chunks — run without a graph. That is fine: for a step of several thousand tokens the launch overhead is a negligible fraction.

With max_num_seqs = 256 the ladder is 1, 2, 4, 8, 16, 24, …, 248, 256, 272, …, 512: 51 sizes. In the default mode each is captured twice (full and piecewise).

--performance-mode interactivity replaces the bottom of the ladder with every size from 1 to 32, so that a server with a handful of concurrent users never pads at all.

What it costs#

CostSizeMitigation
Startup timeSeconds to minutes: compilation on first start, graph capture on every startPersist the compile cache; lower -O level; fewer capture sizes
GPU memoryEach recorded graph keeps its intermediate buffers. Hundreds of megabytes to a few gigabytes in total.Since v0.21.0 this is estimated and subtracted from the KV cache budget (Sizing the Cache); reduce sizes or use FULL_DECODE_ONLY
PaddingA batch between two rungs computes a few extra rowsNegligible above ~16; use interactivity mode for tiny batches

The startup log reports the second and first directly: Graph capturing finished in 9 secs, took 0.31 GiB.

Optimisation levels#

Four presets bundle these choices, selected like a compiler flag:

LevelCompilationCUDA graphsFusionsUse
-O0NoneNONENoneFastest startup; development and debugging
-O1OnPIECEWISEBasicFast startup with the main speedups
-O2On, more compile rangesFULL_AND_PIECEWISEMoreDefault. Production.
-O3Currently identical to -O2Reserved for slower or experimental optimisations
Shell
vllm serve <model> -O1

Any flag you set explicitly overrides the preset. -O0 is the first thing to try when a model fails during startup: if it then works, the problem is in compilation or capture, not in the model.

When it goes wrong#

SymptomLikely causeFirst step
Startup takes many minutes every timeCompile cache not persisted between startsMount the cache directory
Out of memory during startup, after the KV cache is createdGraph capture needs more than was estimatedLower --gpu-memory-utilization, or reduce capture sizes
Wrong output only with graphs onA backend or custom operation is not capture-safeReproduce with -O0, then with -cc.cudagraph_mode=PIECEWISE
Fast at high concurrency, slow at 1 to 3 usersPadding and small-batch overhead--performance-mode interactivity
A stale cache after patching model codeHash did not cover the changeVLLM_DISABLE_COMPILE_CACHE=1 once

Code#

The capture ladder, and the padding it causes.

Go
package main

import "fmt"

// captureSizes reproduces the default ladder in VllmConfig._set_cudagraph_sizes.
func captureSizes(maxSize int, interactivity bool) []int {
	var sizes []int
	if interactivity {
		for i := 1; i <= min(maxSize, 32); i++ {
			sizes = append(sizes, i)
		}
	} else {
		for _, i := range []int{1, 2, 4} {
			if i <= maxSize {
				sizes = append(sizes, i)
			}
		}
	}
	for i := 8; i < min(maxSize+1, 256); i += 8 {
		if !contains(sizes, i) {
			sizes = append(sizes, i)
		}
	}
	for i := 256; i <= maxSize; i += 16 {
		sizes = append(sizes, i)
	}
	return sizes
}

func contains(xs []int, x int) bool {
	for _, v := range xs {
		if v == x {
			return true
		}
	}
	return false
}

// padTo returns the smallest captured size >= n, or 0 if none (eager execution).
func padTo(sizes []int, n int) int {
	best := 0
	for _, s := range sizes {
		if s >= n && (best == 0 || s < best) {
			best = s
		}
	}
	return best
}

func main() {
	const maxNumSeqs = 256
	maxSize := min(maxNumSeqs*1*2, 512) // decode_query_len = 1 without speculative decoding

	def := captureSizes(maxSize, false)
	inter := captureSizes(maxSize, true)
	fmt.Printf("max_num_seqs=%d -> max capture size %d\n", maxNumSeqs, maxSize)
	fmt.Printf("default ladder:        %d sizes: %v ... %v\n", len(def), def[:8], def[len(def)-3:])
	fmt.Printf("interactivity ladder:  %d sizes\n\n", len(inter))

	fmt.Printf("%8s %10s %8s %14s %8s\n", "batch", "default", "waste", "interactivity", "waste")
	for _, n := range []int{1, 3, 5, 9, 17, 33, 100, 250, 300, 600} {
		d, i := padTo(def, n), padTo(inter, n)
		row := func(p int) (string, string) {
			if p == 0 {
				return "eager", "-"
			}
			return fmt.Sprint(p), fmt.Sprintf("%.0f%%", 100*float64(p-n)/float64(p))
		}
		ds, dw := row(d)
		is, iw := row(i)
		fmt.Printf("%8d %10s %8s %14s %8s\n", n, ds, dw, is, iw)
	}

	// Average waste if every batch size from 1 to maxSize were equally likely.
	total := func(sizes []int) float64 {
		w := 0.0
		for n := 1; n <= maxSize; n++ {
			p := padTo(sizes, n)
			w += float64(p-n) / float64(p)
		}
		return 100 * w / float64(maxSize)
	}
	fmt.Printf("\naverage padding over all batch sizes: default %.1f%%, interactivity %.1f%%\n",
		total(def), total(inter))
}

The waste is largest for small batches — nine requests padded to sixteen is 44% — and is below 5% by a hundred. That is the whole case for the interactivity ladder: it spends capture time and memory on twenty-five extra sizes to remove padding exactly where it is proportionally worst.

Remember this#

  • Launch overhead matters most for small models and small batches; CUDA graphs remove it by replaying a recording.
  • A recording is tied to one token count, so vLLM captures a ladder of sizes and pads up.
  • torch.compile is on by default, finishes before any request is served, and is cached on disk.
  • The default mode is FULL_AND_PIECEWISE: whole-pass graphs for pure decode, piecewise graphs otherwise.
  • The attention backend limits which modes are possible; vLLM downgrades automatically.
  • -O0 to -O3 trade startup time for speed; -O2 is the default.
  • Persist ~/.cache/vllm or pay the compilation time on every cold start.

Try it#

  1. In the program, set maxNumSeqs to 64. What is the top of the ladder, and what happens to a batch of 200 tokens?
  2. Start the same model with -O0, -O1 and -O2. Record startup time and tokens per second at 4 concurrent requests. Which level gives the best value for a development loop?
  3. Delete ~/.cache/vllm/torch_compile_cache, start a server, stop it, and start it again. Compare the two init engine … took lines.

Check yourself#

  1. Why can a CUDA graph be reused for different requests but not for a different batch size?
  2. What makes a batch “uniform”, and why does that case get a fully recorded graph?
  3. Why does vLLM compile everything at startup instead of when a new shape first appears?

Sources#

Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom