Pidoku

Chunked Prefill and the Token Budget

Intermediate 50 min Difficulty 3/5 Lesson 02 of 04

Prerequisites One Scheduling Step

The idea in one minute#

--max-num-batched-tokens is the single most consequential tuning flag in vLLM, and it is one number doing two opposite jobs. A request reading its prompt wants the budget large, so the prompt is consumed in few steps and the first token arrives sooner. A request that is already generating wants it small, because it gets one token per step and a step that also reads 8,000 prompt tokens for somebody else takes far longer than a step that does not. Chunked prefill is the mechanism that lets both kinds of request share a batch; the budget is the dial that decides who pays. This lesson shows what the dial is set to by default on your hardware, what moves it, and how to choose a value from the latency you are willing to accept.

A picture#

flowchart LR
  subgraph S1["Step N — budget 8192"]
    direction TB
    D1["32 decoding requests<br/><small>32 tokens</small>"]
    P1["1 prefilling request<br/><small>8160 tokens of a 30,000-token prompt</small>"]
  end
  subgraph S2["Step N+1"]
    direction TB
    D2["32 decoding requests<br/><small>32 tokens</small>"]
    P2["same request<br/><small>next 8160 tokens</small>"]
  end
  S1 --> S2
  S2 --> T[":i-clock: <b>Every decoding request waits<br/>for the whole step</b><br/><small>its inter-token latency = step time</small>"]
  class D1,D2 compute
  class P1,P2 queue
  class T warn

How it really works#

Why mix prefill and decode at all#

The two halves of a request stress different parts of the GPU.

  • Prefill processes many tokens at once. It is limited by arithmetic: the GPU’s compute units are busy.
  • Decode processes one token per request. It is limited by memory bandwidth: the GPU spends its time reading weights, and its compute units are mostly idle (Prefill vs Decode).

A batch that contains both uses the idle compute of the decode work for the prefill work. The project’s documentation gives exactly this reason: chunked prefill “helps achieve better GPU utilization by locating compute-bound (prefill) and memory-bound (decode) requests to the same batch.”

Without chunking, a long prompt has to go into the model in one piece, and a forward pass cannot be interrupted. Every other request would freeze until it finished. With chunking, the prompt is spread over several steps and the decoders get a token in each one.

Where chunking happens in the code#

Nowhere in particular, which is the elegant part. From One Scheduling Step:

Python
num_new_tokens = request.num_tokens - num_computed_tokens
...
num_new_tokens = min(num_new_tokens, request_token_budget)

A prompt larger than the remaining budget is clipped to the budget; num_computed_tokens advances by that amount; next step the request is in running with a smaller gap. Because running requests are scheduled in admission order and decoders need one token each, the decoders always take their share first and the prefill receives what is left.

The only place enable_chunked_prefill is consulted is to forbid clipping:

Python
if (
    not self.scheduler_config.enable_chunked_prefill
    and num_new_tokens > request_token_budget
    and not self.is_mm_encoder_only
):
    # If chunked_prefill is disabled, we can stop the scheduling here.
    break

With chunking off, a prompt that does not fit in the remaining budget waits for a step where it fits whole. Consequently the budget must be at least max_model_len, or long prompts could never run; vLLM raises it automatically in that case and refuses to start if you set it lower.

Chunked prefill is on by default and is switched off automatically for:

  • encoder-decoder models (Whisper-style),
  • models with no KV cache (encoder-only embedding models),
  • models with a non-causal attention layer, where a token’s result depends on tokens after it.

What the budget defaults to#

You almost never see the hard-coded default of 2,048 from SchedulerConfig. vllm serve picks a value from the GPU’s memory, in EngineArgs.get_batch_defaults:

Devicemax_num_batched_tokens (server)(offline LLM)max_num_seqs
GPU with ≥ 160 GiB (B200-class)16,38416,3841,024
GPU with ≥ 70 GiB, not an A100 (H100, H200)8,19216,3841,024
Any other GPU (A100, L4, A10G, RTX…)2,0488,192256
CPU2,048 × world size4,096 × world size128 × world size (server)

The A100 exception has a comment explaining it: “Setting large max_num_batched_tokens for A100 reduces throughput”. An 80 GiB A100 is therefore treated like a small GPU.

Then three adjustments are applied in order:

  1. --performance-mode throughput doubles both numbers (if you did not set them).
  2. The token budget is capped at max_num_seqs × max_model_len.
  3. max_num_seqs is capped at the token budget: there is no point allowing more sequences than tokens.

The value in force is printed at startup: Chunked prefill is enabled with max_num_batched_tokens=8192.

note

The offline LLM class gets larger budgets than the server because a batch job has no clients watching tokens arrive. If you benchmark with the LLM class and deploy with vllm serve, you are measuring a different configuration.

The trade-off, with numbers#

Let a decode step for the current batch take d milliseconds, and let prefill cost p milliseconds per 1,000 tokens. A step that carries c prefill tokens takes roughly:

step time ≈ d + p × c / 1000

Say d = 12 ms and p = 20 ms (an 8B-class model on a large GPU; measure your own). Then for a 32,000-token prompt arriving while others are generating:

BudgetChunksLongest step during the prefillTime to first token for the long promptInter-token latency for everyone else
2,0481612 + 41 = 53 ms0.83 s53 ms
8,192412 + 164 = 176 ms0.69 s176 ms
16,384212 + 328 = 340 ms0.66 s340 ms
32,768112 + 640 = 652 ms0.65 s652 ms — a visible freeze

Two things stand out. The long prompt’s time to first token barely changes: the total prefill work is the same, and only the fixed d per chunk is saved. But every other request’s inter-token latency scales directly with the budget. Raising the budget buys a little TTFT and total throughput at the price of smooth streaming.

That is why the project’s guidance reads:

  • “Smaller values (e.g., 2048) achieve better ITL because there are fewer prefills slowing down decodes.”
  • “Higher values achieve better time to first token (TTFT).”
  • “For optimal throughput, we recommend setting max_num_batched_tokens > 8192 especially for smaller models on large GPUs.”

Choosing a value#

Work backwards from the inter-token latency you can tolerate.

tolerable step time   T    e.g. 80 ms for a chat product
decode step time      d    measure with only decoders running
prefill cost          p    ms per 1,000 tokens, measure with one long prompt

budget ≈ (T − d) / p × 1000
       = (80 − 12) / 20 × 1000  ≈ 3,400 tokens

For a batch or agent workload with no human reading tokens as they arrive, take the largest value that does not reduce throughput. For a mixed workload, the next section’s flag is usually the better tool.

--long-prefill-token-threshold: a cap per request#

The budget is shared by all prefilling requests in a step. One 100,000-token prompt can take the entire budget step after step, and ten 500-token prompts that arrived just after it wait behind it. A per-request cap fixes that:

Python
if 0 < long_prefill_token_threshold < num_new_tokens:
    num_new_tokens = long_prefill_token_threshold

With --long-prefill-token-threshold 2048 and a budget of 8,192, no single request takes more than 2,048 tokens of a step. A long prompt and three short ones can all make progress in the same step. The default is 0, meaning no cap.

Two refinements in the source are worth knowing:

Python
# `long_prefill_token_threshold` exists to stop a long prefill from
# starving other requests of the token budget. When it is the only
# request there is nobody to starve, so let it use the whole budget.
long_prefill_token_threshold = (
    self.scheduler_config.long_prefill_token_threshold
    if num_eligible_reqs > 1
    else 0
)

A lone request is never throttled. And on main (not yet in v0.30.0) there is --long-prefill-token-threshold-adaptive, which raises the cap to a fair share of the budget when few requests are present:

Python
long_prefill_token_threshold = max(
    long_prefill_token_threshold, input_budget // num_eligible_reqs
)

With a budget of 8,192 and two requests, each may take 4,096 regardless of a lower configured threshold, so the budget is not left unused.

FlagDefaultWhat it does
--max-num-seqsBy GPU, see aboveMost requests in running. It also sizes per-request buffers in the model runner and the largest CUDA graph, so it costs memory and startup time.
--max-num-active-seqs (main only)unsetAdmission limit below max_num_seqs: keeps decode batches smaller without shrinking buffers or graphs.
--scheduler-reserve-full-islonAdmit a request only if its whole prompt fits in free KV blocks.
--watermark0.0Fraction of KV blocks kept free when admitting. Headroom so running requests can grow without preempting each other.
--disable-chunked-mm-inputoffNever split one image’s placeholder tokens across two steps; the whole image goes in one chunk.
--prefill-schedule-interval1Under data parallelism, admit new prefills only every Nth step, aligned across ranks, so their step times match.
--performance-modebalancedthroughput doubles the two batch limits and prefers larger CUDA graphs and throughput-oriented kernels; interactivity captures a CUDA graph for every batch size from 1 to 32 to remove padding at low concurrency.

What “prioritises decode” does and does not mean#

The scheduler serves running in admission order. A request still in its prompt and a request generating are both in running, and neither is preferred because of its phase. What protects decoders is arithmetic: each needs one token, so even hundreds of them consume a negligible share of the budget before a prefilling request is reached. A prefilling request that was admitted earlier than a decoder is still scheduled before it — but there is always enough budget left for the decoder’s single token unless the budget is very small.

The one case where a decoder is skipped is deliberate: if the remaining budget is exhausted by requests ahead of it. With a budget of 2,048 and 3,000 running decoders that could happen, which is why max_num_seqs is capped at the budget.

Code#

The table above, computed, plus the fairness problem that long_prefill_token_threshold solves: one huge prompt and several small ones arriving together.

Go
package main

import "fmt"

const (
	decodeMs  = 12.0 // a step with only decoders
	prefillMs = 20.0 // per 1,000 prefill tokens
)

type req struct {
	name     string
	prompt   int
	computed int
	ttft     float64
}

// simulate runs prefill for all requests and returns each one's time to first
// token and the slowest step (which is everyone else's worst inter-token latency).
func simulate(budget, perRequestCap int, prompts []int) ([]float64, float64) {
	reqs := make([]*req, len(prompts))
	for i, p := range prompts {
		reqs[i] = &req{name: fmt.Sprint(i), prompt: p}
	}
	clock, worst := 0.0, 0.0
	for remaining := len(reqs); remaining > 0; {
		left, used := budget, 0
		var done []*req
		for _, r := range reqs {
			if r.computed == r.prompt || left == 0 {
				continue
			}
			n := min(r.prompt-r.computed, left)
			if perRequestCap > 0 && remaining > 1 {
				n = min(n, perRequestCap)
			}
			r.computed += n
			left -= n
			used += n
			if r.computed == r.prompt {
				done = append(done, r)
			}
		}
		step := decodeMs + prefillMs*float64(used)/1000
		clock += step
		worst = max(worst, step)
		for _, r := range done {
			r.ttft = clock
			remaining--
		}
	}
	out := make([]float64, len(reqs))
	for i, r := range reqs {
		out[i] = r.ttft
	}
	return out, worst
}

func main() {
	fmt.Println("One 32,000-token prompt, by budget:")
	for _, b := range []int{2048, 8192, 16384, 32768} {
		ttft, worst := simulate(b, 0, []int{32000})
		fmt.Printf("  budget %6d  TTFT %5.2f s   worst step %4.0f ms\n", b, ttft[0]/1000, worst)
	}

	fmt.Println("\nOne 60,000-token prompt followed by four 500-token prompts, budget 8192:")
	prompts := []int{60000, 500, 500, 500, 500}
	for _, cap := range []int{0, 2048} {
		ttft, worst := simulate(8192, cap, prompts)
		fmt.Printf("  long_prefill_token_threshold=%-5d long TTFT %5.2f s   short TTFT %5.2f s   worst step %4.0f ms\n",
			cap, ttft[0]/1000, ttft[1]/1000, worst)
	}
}

In the second experiment the short prompts wait more than a second for their first token without the cap, and a fraction of that with it, while the long prompt loses very little. The cap is close to free and is the first thing to set on a server that mixes long documents with short chats.

Remember this#

  • Chunked prefill is not a separate code path: a prompt larger than the budget is clipped and continues next step.
  • Every request in a batch waits for the whole step, so the budget sets the worst inter-token latency.
  • Server defaults: 2,048 on most GPUs, 8,192 on H100/H200-class, 16,384 at 160 GiB and above. The offline LLM class uses larger values.
  • A bigger budget improves throughput and slightly improves TTFT; it worsens inter-token latency in proportion.
  • --long-prefill-token-threshold caps one request’s share per step and protects short prompts from long ones.
  • Size the budget from the step time you can tolerate: (T − d) / p × 1000.

Try it#

  1. In the program, set prefillMs to 60 (a much larger model). Which budget keeps the worst step under 100 ms?
  2. Run the second experiment with a cap of 512 and of 4,096. Plot (by hand) long-prompt TTFT against short-prompt TTFT. Where is the knee?
  3. On a real server, send one request with a very long prompt while a second client streams a long answer. Measure the second client’s inter-token gaps with the Go client from Your First Server. Repeat with --max-num-batched-tokens 1024.

Check yourself#

  1. Why does raising the budget improve a long prompt’s time to first token only slightly?
  2. What is the default token budget for vllm serve on an A100, and why is it not the same as on an H100?
  3. A server has one request in it. Does --long-prefill-token-threshold limit that request?

Sources#

Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom