The idea in one minute#
--max-num-batched-tokens is the single most consequential tuning flag in vLLM, and it is one
number doing two opposite jobs. A request reading its prompt wants the budget large, so the
prompt is consumed in few steps and the first token arrives sooner. A request that is already
generating wants it small, because it gets one token per step and a step that also reads
8,000 prompt tokens for somebody else takes far longer than a step that does not. Chunked
prefill is the mechanism that lets both kinds of request share a batch; the budget is the dial
that decides who pays. This lesson shows what the dial is set to by default on your hardware,
what moves it, and how to choose a value from the latency you are willing to accept.
A picture#
flowchart LR
subgraph S1["Step N — budget 8192"]
direction TB
D1["32 decoding requests<br/><small>32 tokens</small>"]
P1["1 prefilling request<br/><small>8160 tokens of a 30,000-token prompt</small>"]
end
subgraph S2["Step N+1"]
direction TB
D2["32 decoding requests<br/><small>32 tokens</small>"]
P2["same request<br/><small>next 8160 tokens</small>"]
end
S1 --> S2
S2 --> T[":i-clock: <b>Every decoding request waits<br/>for the whole step</b><br/><small>its inter-token latency = step time</small>"]
class D1,D2 compute
class P1,P2 queue
class T warnHow it really works#
Why mix prefill and decode at all#
The two halves of a request stress different parts of the GPU.
- Prefill processes many tokens at once. It is limited by arithmetic: the GPU’s compute units are busy.
- Decode processes one token per request. It is limited by memory bandwidth: the GPU spends its time reading weights, and its compute units are mostly idle (Prefill vs Decode).
A batch that contains both uses the idle compute of the decode work for the prefill work. The project’s documentation gives exactly this reason: chunked prefill “helps achieve better GPU utilization by locating compute-bound (prefill) and memory-bound (decode) requests to the same batch.”
Without chunking, a long prompt has to go into the model in one piece, and a forward pass cannot be interrupted. Every other request would freeze until it finished. With chunking, the prompt is spread over several steps and the decoders get a token in each one.
Where chunking happens in the code#
Nowhere in particular, which is the elegant part. From One Scheduling Step:
num_new_tokens = request.num_tokens - num_computed_tokens
...
num_new_tokens = min(num_new_tokens, request_token_budget)A prompt larger than the remaining budget is clipped to the budget; num_computed_tokens
advances by that amount; next step the request is in running with a smaller gap. Because
running requests are scheduled in admission order and decoders need one token each, the
decoders always take their share first and the prefill receives what is left.
The only place enable_chunked_prefill is consulted is to forbid clipping:
if (
not self.scheduler_config.enable_chunked_prefill
and num_new_tokens > request_token_budget
and not self.is_mm_encoder_only
):
# If chunked_prefill is disabled, we can stop the scheduling here.
breakWith chunking off, a prompt that does not fit in the remaining budget waits for a step where it
fits whole. Consequently the budget must be at least max_model_len, or long prompts could
never run; vLLM raises it automatically in that case and refuses to start if you set it lower.
Chunked prefill is on by default and is switched off automatically for:
- encoder-decoder models (Whisper-style),
- models with no KV cache (encoder-only embedding models),
- models with a non-causal attention layer, where a token’s result depends on tokens after it.
What the budget defaults to#
You almost never see the hard-coded default of 2,048 from SchedulerConfig. vllm serve
picks a value from the GPU’s memory, in EngineArgs.get_batch_defaults:
| Device | max_num_batched_tokens (server) | (offline LLM) | max_num_seqs |
|---|---|---|---|
| GPU with ≥ 160 GiB (B200-class) | 16,384 | 16,384 | 1,024 |
| GPU with ≥ 70 GiB, not an A100 (H100, H200) | 8,192 | 16,384 | 1,024 |
| Any other GPU (A100, L4, A10G, RTX…) | 2,048 | 8,192 | 256 |
| CPU | 2,048 × world size | 4,096 × world size | 128 × world size (server) |
The A100 exception has a comment explaining it: “Setting large max_num_batched_tokens for
A100 reduces throughput”. An 80 GiB A100 is therefore treated like a small GPU.
Then three adjustments are applied in order:
--performance-mode throughputdoubles both numbers (if you did not set them).- The token budget is capped at
max_num_seqs × max_model_len. max_num_seqsis capped at the token budget: there is no point allowing more sequences than tokens.
The value in force is printed at startup: Chunked prefill is enabled with max_num_batched_tokens=8192.
note
The offline LLM class gets larger budgets than the server because a batch job has no
clients watching tokens arrive. If you benchmark with the LLM class and deploy with
vllm serve, you are measuring a different configuration.
The trade-off, with numbers#
Let a decode step for the current batch take d milliseconds, and let prefill cost p
milliseconds per 1,000 tokens. A step that carries c prefill tokens takes roughly:
step time ≈ d + p × c / 1000Say d = 12 ms and p = 20 ms (an 8B-class model on a large GPU; measure your own). Then for
a 32,000-token prompt arriving while others are generating:
| Budget | Chunks | Longest step during the prefill | Time to first token for the long prompt | Inter-token latency for everyone else |
|---|---|---|---|---|
| 2,048 | 16 | 12 + 41 = 53 ms | 0.83 s | 53 ms |
| 8,192 | 4 | 12 + 164 = 176 ms | 0.69 s | 176 ms |
| 16,384 | 2 | 12 + 328 = 340 ms | 0.66 s | 340 ms |
| 32,768 | 1 | 12 + 640 = 652 ms | 0.65 s | 652 ms — a visible freeze |
Two things stand out. The long prompt’s time to first token barely changes: the total prefill
work is the same, and only the fixed d per chunk is saved. But every other request’s
inter-token latency scales directly with the budget. Raising the budget buys a little TTFT and
total throughput at the price of smooth streaming.
That is why the project’s guidance reads:
- “Smaller values (e.g., 2048) achieve better ITL because there are fewer prefills slowing down decodes.”
- “Higher values achieve better time to first token (TTFT).”
- “For optimal throughput, we recommend setting
max_num_batched_tokens > 8192especially for smaller models on large GPUs.”
Choosing a value#
Work backwards from the inter-token latency you can tolerate.
tolerable step time T e.g. 80 ms for a chat product
decode step time d measure with only decoders running
prefill cost p ms per 1,000 tokens, measure with one long prompt
budget ≈ (T − d) / p × 1000
= (80 − 12) / 20 × 1000 ≈ 3,400 tokensFor a batch or agent workload with no human reading tokens as they arrive, take the largest value that does not reduce throughput. For a mixed workload, the next section’s flag is usually the better tool.
--long-prefill-token-threshold: a cap per request#
The budget is shared by all prefilling requests in a step. One 100,000-token prompt can take the entire budget step after step, and ten 500-token prompts that arrived just after it wait behind it. A per-request cap fixes that:
if 0 < long_prefill_token_threshold < num_new_tokens:
num_new_tokens = long_prefill_token_thresholdWith --long-prefill-token-threshold 2048 and a budget of 8,192, no single request takes more
than 2,048 tokens of a step. A long prompt and three short ones can all make progress in the
same step. The default is 0, meaning no cap.
Two refinements in the source are worth knowing:
# `long_prefill_token_threshold` exists to stop a long prefill from
# starving other requests of the token budget. When it is the only
# request there is nobody to starve, so let it use the whole budget.
long_prefill_token_threshold = (
self.scheduler_config.long_prefill_token_threshold
if num_eligible_reqs > 1
else 0
)A lone request is never throttled. And on main (not yet in v0.30.0) there is
--long-prefill-token-threshold-adaptive, which raises the cap to a fair share of the budget
when few requests are present:
long_prefill_token_threshold = max(
long_prefill_token_threshold, input_budget // num_eligible_reqs
)With a budget of 8,192 and two requests, each may take 4,096 regardless of a lower configured threshold, so the budget is not left unused.
Related controls#
| Flag | Default | What it does |
|---|---|---|
--max-num-seqs | By GPU, see above | Most requests in running. It also sizes per-request buffers in the model runner and the largest CUDA graph, so it costs memory and startup time. |
--max-num-active-seqs (main only) | unset | Admission limit below max_num_seqs: keeps decode batches smaller without shrinking buffers or graphs. |
--scheduler-reserve-full-isl | on | Admit a request only if its whole prompt fits in free KV blocks. |
--watermark | 0.0 | Fraction of KV blocks kept free when admitting. Headroom so running requests can grow without preempting each other. |
--disable-chunked-mm-input | off | Never split one image’s placeholder tokens across two steps; the whole image goes in one chunk. |
--prefill-schedule-interval | 1 | Under data parallelism, admit new prefills only every Nth step, aligned across ranks, so their step times match. |
--performance-mode | balanced | throughput doubles the two batch limits and prefers larger CUDA graphs and throughput-oriented kernels; interactivity captures a CUDA graph for every batch size from 1 to 32 to remove padding at low concurrency. |
What “prioritises decode” does and does not mean#
The scheduler serves running in admission order. A request still in its prompt and a request
generating are both in running, and neither is preferred because of its phase. What protects
decoders is arithmetic: each needs one token, so even hundreds of them consume a negligible
share of the budget before a prefilling request is reached. A prefilling request that was
admitted earlier than a decoder is still scheduled before it — but there is always enough
budget left for the decoder’s single token unless the budget is very small.
The one case where a decoder is skipped is deliberate: if the remaining budget is exhausted by
requests ahead of it. With a budget of 2,048 and 3,000 running decoders that could happen, which
is why max_num_seqs is capped at the budget.
Code#
The table above, computed, plus the fairness problem that long_prefill_token_threshold
solves: one huge prompt and several small ones arriving together.
package main
import "fmt"
const (
decodeMs = 12.0 // a step with only decoders
prefillMs = 20.0 // per 1,000 prefill tokens
)
type req struct {
name string
prompt int
computed int
ttft float64
}
// simulate runs prefill for all requests and returns each one's time to first
// token and the slowest step (which is everyone else's worst inter-token latency).
func simulate(budget, perRequestCap int, prompts []int) ([]float64, float64) {
reqs := make([]*req, len(prompts))
for i, p := range prompts {
reqs[i] = &req{name: fmt.Sprint(i), prompt: p}
}
clock, worst := 0.0, 0.0
for remaining := len(reqs); remaining > 0; {
left, used := budget, 0
var done []*req
for _, r := range reqs {
if r.computed == r.prompt || left == 0 {
continue
}
n := min(r.prompt-r.computed, left)
if perRequestCap > 0 && remaining > 1 {
n = min(n, perRequestCap)
}
r.computed += n
left -= n
used += n
if r.computed == r.prompt {
done = append(done, r)
}
}
step := decodeMs + prefillMs*float64(used)/1000
clock += step
worst = max(worst, step)
for _, r := range done {
r.ttft = clock
remaining--
}
}
out := make([]float64, len(reqs))
for i, r := range reqs {
out[i] = r.ttft
}
return out, worst
}
func main() {
fmt.Println("One 32,000-token prompt, by budget:")
for _, b := range []int{2048, 8192, 16384, 32768} {
ttft, worst := simulate(b, 0, []int{32000})
fmt.Printf(" budget %6d TTFT %5.2f s worst step %4.0f ms\n", b, ttft[0]/1000, worst)
}
fmt.Println("\nOne 60,000-token prompt followed by four 500-token prompts, budget 8192:")
prompts := []int{60000, 500, 500, 500, 500}
for _, cap := range []int{0, 2048} {
ttft, worst := simulate(8192, cap, prompts)
fmt.Printf(" long_prefill_token_threshold=%-5d long TTFT %5.2f s short TTFT %5.2f s worst step %4.0f ms\n",
cap, ttft[0]/1000, ttft[1]/1000, worst)
}
}In the second experiment the short prompts wait more than a second for their first token without the cap, and a fraction of that with it, while the long prompt loses very little. The cap is close to free and is the first thing to set on a server that mixes long documents with short chats.
Remember this#
- Chunked prefill is not a separate code path: a prompt larger than the budget is clipped and continues next step.
- Every request in a batch waits for the whole step, so the budget sets the worst inter-token latency.
- Server defaults: 2,048 on most GPUs, 8,192 on H100/H200-class, 16,384 at 160 GiB and above. The offline
LLMclass uses larger values. - A bigger budget improves throughput and slightly improves TTFT; it worsens inter-token latency in proportion.
--long-prefill-token-thresholdcaps one request’s share per step and protects short prompts from long ones.- Size the budget from the step time you can tolerate:
(T − d) / p × 1000.
Try it#
- In the program, set
prefillMsto 60 (a much larger model). Which budget keeps the worst step under 100 ms? - Run the second experiment with a cap of 512 and of 4,096. Plot (by hand) long-prompt TTFT against short-prompt TTFT. Where is the knee?
- On a real server, send one request with a very long prompt while a second client streams a
long answer. Measure the second client’s inter-token gaps with the Go client from
Your First Server. Repeat with
--max-num-batched-tokens 1024.
Check yourself#
- Why does raising the budget improve a long prompt’s time to first token only slightly?
- What is the default token budget for
vllm serveon an A100, and why is it not the same as on an H100? - A server has one request in it. Does
--long-prefill-token-thresholdlimit that request?
Sources#
Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.
vllm/v1/core/sched/scheduler.pyvllm/engine/arg_utils.py— the hardware-dependent defaultsvllm/config/scheduler.py- Optimization and tuning: chunked prefill