The idea in one minute#
vLLM has hundreds of flags and perhaps a dozen that matter. Tuning is not trying them at random; it is deciding what you are optimising for, measuring with a load that looks like yours, and changing the one setting that controls the mechanism you found to be the bottleneck. The earlier lessons gave you the mechanisms. This one gives the measurement method — including the single most common benchmarking mistake, which makes an overloaded server look healthy — and a table from symptom to flag.
A picture#
flowchart LR G[":i-scale: <b>1. Choose the goal</b><br/><small>latency SLO, throughput, or cost</small>"] --> W[":i-file-text: <b>2. Describe the workload</b><br/><small>prompt and output length distributions,<br/>arrival rate, shared prefixes</small>"] W --> B[":i-gauge: <b>3. Measure</b><br/><small>open-loop load at several rates</small>"] B --> K[":i-search: <b>4. Find the bottleneck</b><br/><small>metrics: queue, cache, step time, CPU</small>"] K --> F[":i-wrench: <b>5. Change one setting</b>"] F --> B class G,W neutral class B compute class K queue class F warn
How it really works#
Three goals that pull apart#
| Goal | You care about | Typical workload | Direction of the main knobs |
|---|---|---|---|
| Interactivity | Time to first token, even streaming | Chat, coding assistants | Smaller token budget, fine-grained CUDA graphs, headroom in the cache |
| Throughput | Tokens per second per GPU | Batch jobs, offline evaluation, agents | Larger token budget and max_num_seqs, full cache |
| Cost under an SLO | Requests per GPU that still meet a latency target | Most production services | Tune for throughput until the SLO breaks, then back off |
vLLM exposes the first two directly:
vllm serve <model> --performance-mode interactivity # or: throughput, balanced (default)throughput doubles max_num_batched_tokens and max_num_seqs and prefers larger CUDA graphs
and throughput-oriented kernels. interactivity captures a CUDA graph for every batch size
from 1 to 32 and prefers latency-oriented kernels
(CUDA Graphs and Compilation). Start
from the mode that matches your goal and adjust from there.
The flags that matter#
| Flag | Mechanism | Raise it to… | Cost of raising it |
|---|---|---|---|
--max-num-batched-tokens | Tokens per step | Speed up prefill, raise throughput | Longer steps: worse inter-token latency; less KV cache |
--max-num-seqs | Running requests per engine | Admit more concurrency | Bigger batches, more preemption risk, more graph memory |
--long-prefill-token-threshold | Per-request share of a step | (lower it) to protect short prompts | Long prompts take more steps |
--gpu-memory-utilization | Size of the block pool | Add cache capacity | Less safety margin |
--max-model-len | Longest request | Accept longer requests | Lower worst-case concurrency; may not start |
--kv-cache-dtype fp8 | Bytes per cached token | (set it) to double cache capacity | Small quality risk |
--watermark | Admission headroom | Reduce preemption | Slightly lower concurrency |
--max-num-queued-reqs, --max-num-queued-tokens | Admission control | (set them) to bound latency under overload | Some requests get 503 |
--stream-interval | Tokens per streamed event | Cut per-token CPU overhead | Chunkier streaming |
--api-server-count | Frontend processes | Remove a tokenising or media bottleneck | CPU cores |
-O, --performance-mode | Compilation and graph ladder | — | Startup time and memory |
--tensor-parallel-size, --data-parallel-size | GPU layout | — | See Parallelism |
Measuring: the built-in tools#
vllm bench serve \
--model <model> --endpoint /v1/chat/completions \
--dataset-name random --random-input-len 1024 --random-output-len 256 \
--num-prompts 2000 --request-rate 20It reports, measured at the client:
Request throughput (req/s): 1.73
Output token throughput (tok/s): 382.89
Mean TTFT (ms): 71.54 P99 TTFT (ms): 79.49
Mean TPOT (ms): 7.91 P99 TPOT (ms): 8.03
Mean ITL (ms): 7.74 P99 ITL (ms): 8.39| Command | Use |
|---|---|
vllm bench serve | Load a running server. Datasets: random, ShareGPT, long-document, multimodal, a prefix-caching workload, timed trace replay. |
vllm bench throughput | Offline: the LLM class with no server. An upper bound, not a prediction of serving performance. |
vllm bench latency | Offline: one request at a time. |
vllm bench sweep serve | Start the server and run bench serve across a grid of server and client parameters, resetting caches between runs. |
For production-grade load testing the documentation points to GuideLLM, a separate project under the same organisation.
Three load-shaping options decide whether your results mean anything:
--request-rate R: an open-loop load. Requests arrive atRper second regardless of how the server is doing.infsends everything at once.--max-concurrency C: a closed-loop limit. At mostCrequests in flight; a new one starts only when one finishes.--burstiness: 1.0 gives Poisson arrivals; lower values are burstier.
The mistake: closed-loop benchmarks hide overload#
With only --max-concurrency 64, sixty-four clients each send a request, wait for the answer,
and send the next. If the server slows down, the clients slow down with it. The arrival rate
adapts to the server. No queue can ever build beyond 64, so latency looks stable at any
“load” — because you are not controlling load at all.
Real users do not wait for each other. They arrive at their own rate. If that rate exceeds what the server can do, the queue grows without bound and time to first token climbs every second. A closed-loop test cannot show that. It is called coordinated omission, and it is the reason teams ship servers that “benchmarked fine at 64 concurrency” and fall over in production.
The fix: drive load with --request-rate at several values and plot latency against rate. The
curve is flat, then bends sharply. The rate just before the bend is the server’s capacity for
your SLO. --max-concurrency may be added as a safety cap, but the rate must be the
independent variable.
Other ways to fool yourself#
- Warm caches. The documentation warns: “Repeating
vllm bench serveagainst the same server can reuse prompts left in the prefix cache and inflate throughput.” Vary--seed, restart, or use the sweep tool, which resets caches. - The wrong length distribution. Fixed 1,024-in, 256-out traffic has no long prompts to disturb decoders and no variance to cause preemption. Use lengths sampled from your logs: the request-shape histograms in Metrics give them.
- No shared prefixes, or too many. A random dataset has a 0% prefix-cache hit rate; a dataset that repeats one prompt has 100%. Neither is your workload.
- Offline numbers for an online service. The
LLMclass gets a larger token budget by default and has no HTTP, tokenising-under-load or streaming cost. - Measuring the client. A Python load generator on one core saturates before a fast server does. Watch the client’s CPU.
- Averages. Mean latency hides the tail that users complain about. Report p50, p95, p99.
- Short runs. The first seconds include queue build-up and cold caches. Run long enough to
reach steady state, and check
num_requests_waitingis not still rising when you stop.
From symptom to setting#
| Symptom (with the metric) | Mechanism | Change |
|---|---|---|
High TTFT, high request_queue_time, running = max_num_seqs | Sequence limit | Raise --max-num-seqs |
High TTFT, high queue time, kv_cache_usage_perc ≈ 1 | Memory | FP8 cache, quantised weights, lower --max-model-len, more GPUs |
num_preemptions rising | Over-admission | Lower --max-num-seqs, set --watermark 0.05 |
| High TTFT, low queue time, long prompts | Prefill speed | Raise --max-num-batched-tokens; fix prompt order for cache hits |
| High TTFT, low queue and prefill time | Frontend CPU | More --api-server-count, more cores; check media loading |
| p99 inter-token latency ≫ p50 | Prefill chunks in decode steps | Lower --max-num-batched-tokens; --long-prefill-token-threshold; disaggregate |
| Short prompts wait behind long ones | Budget monopolised | --long-prefill-token-threshold 2048 |
| All inter-token latency high, GPU not full | CPU-bound engine core | Dedicated physical core; --stream-interval; fewer structured-output requests per engine |
| Slow at 1 to 4 users, fine at 64 | Padding and launch overhead | --performance-mode interactivity |
| Latency grows without bound under load | No admission control | --max-num-queued-reqs or --max-num-queued-tokens |
| Throughput flat while adding concurrency | Compute-saturated | Add replicas; quantise; this GPU is full |
| Startup takes minutes | Compile and capture | Persist ~/.cache/vllm; -O1; vllm preload |
A tuning session, in order#
- Fit. Make the model start with the
max_model_lenyour traffic needs. Read the capacity line; checkcapacity ÷ p95 request tokenscomfortably exceedsmax_num_seqs. - Baseline. Default flags, your length distribution, open-loop load at a low rate. Record p50/p95/p99 TTFT and inter-token latency.
- Find the knee. Raise the rate in steps until p95 TTFT or p99 inter-token latency breaks your SLO. Note what the metrics say is binding there.
- Move the binding constraint. One flag, from the table. Re-run step 3.
- Stop when the constraint is the GPU’s compute, or when the remaining changes trade one SLO for another.
- Protect it. Set admission limits a little above the knee so that overload produces fast
503s instead of slow timeouts, and configure the autoscaler on
num_requests_waitingor queue time rather than GPU utilisation, which is high whenever the server is busy at all.
Code#
Why a closed-loop test cannot find the knee. The same server is measured two ways: a fixed number of clients, and a fixed arrival rate.
package main
import "fmt"
// The server: it can work on up to maxRunning requests at once, each needing
// serviceSeconds of attention when running alone; with n running, each
// progresses at min(1, capacity/n) of full speed.
const (
maxRunning = 64
capacity = 32.0 // requests that can run at full speed simultaneously
serviceSeconds = 2.0
dt = 0.01
duration = 300.0
)
type req struct {
arrived, started float64
left float64
}
func simulate(arrivalRate float64, closedClients int) (throughput, meanLatency, worstQueueWait float64) {
var waiting, running []*req
done, totalLatency, carry := 0, 0.0, 0.0
if closedClients > 0 {
for i := 0; i < closedClients; i++ {
waiting = append(waiting, &req{left: serviceSeconds})
}
}
for t := 0.0; t < duration; t += dt {
if closedClients == 0 { // open loop: arrivals do not care how the server is doing
carry += arrivalRate * dt
for ; carry >= 1; carry-- {
waiting = append(waiting, &req{arrived: t, left: serviceSeconds})
}
}
for len(waiting) > 0 && len(running) < maxRunning {
r := waiting[0]
waiting = waiting[1:]
r.started = t
worstQueueWait = max(worstQueueWait, t-r.arrived)
running = append(running, r)
}
speed := min(1.0, capacity/float64(max(len(running), 1)))
keep := running[:0]
for _, r := range running {
if r.left -= speed * dt; r.left > 0 {
keep = append(keep, r)
continue
}
done++
totalLatency += t - r.arrived
if closedClients > 0 { // closed loop: the client sends its next request only now
waiting = append(waiting, &req{arrived: t, left: serviceSeconds})
}
}
running = keep
}
return float64(done) / duration, totalLatency / float64(max(done, 1)), worstQueueWait
}
func main() {
fmt.Printf("server capacity: %.0f requests/s at full speed\n\n", capacity/serviceSeconds)
fmt.Println("closed loop (--max-concurrency only)")
fmt.Printf(" %8s %12s %14s %18s\n", "clients", "req/s", "mean latency", "worst queue wait")
for _, c := range []int{8, 32, 64, 128, 256} {
tp, lat, q := simulate(0, c)
fmt.Printf(" %8d %12.1f %12.2f s %16.2f s\n", c, tp, lat, q)
}
fmt.Println("\nopen loop (--request-rate)")
fmt.Printf(" %8s %12s %14s %18s\n", "rate", "req/s", "mean latency", "worst queue wait")
for _, r := range []float64{4, 8, 12, 15, 17, 20, 24} {
tp, lat, q := simulate(r, 0)
fmt.Printf(" %8.0f %12.1f %12.2f s %16.2f s\n", r, tp, lat, q)
}
}In the closed-loop table, latency grows with the number of clients but every row is a steady, finite number, and throughput sits at capacity: 256 clients look like a slower but stable server. In the open-loop table, everything is fine up to the capacity and then the worst queue wait grows with the length of the test, because the queue never stops growing. Only the second table tells you where the cliff is.
Remember this#
- Decide the goal first: interactivity, throughput, or cost under an SLO.
--performance-modesets a starting point. - About a dozen flags matter, and each controls one mechanism from earlier lessons.
- Benchmark with open-loop load (
--request-rate) at several rates; closed-loop tests hide overload. - Use your real length distribution and prefix-sharing pattern; reset or vary to avoid warm-cache inflation.
- Report tail percentiles, not means.
- Use the metrics to name the binding constraint, change one flag, and re-measure.
- Set admission limits just above the knee, and autoscale on queue signals, not GPU utilisation.
Try it#
- In the program, raise
maxRunningto 128 with the samecapacity. What happens to open-loop latency just below the knee? This is the cost of amax_num_seqsthat is too high for the hardware. - Add admission control to the open-loop case: reject an arrival when more than 40 are waiting. Report the rejection rate and the worst queue wait at a rate of 20.
- Run
vllm bench serveagainst a real server at five request rates and plot p95 TTFT. Mark your SLO and read off the capacity.
Check yourself#
- Why does a benchmark using only
--max-concurrencyfail to reveal an overloaded server? - p99 inter-token latency is six times the median while GPU utilisation is modest. Which mechanism is the likely cause and which two flags address it?
- Why is GPU utilisation a poor autoscaling signal for an inference server?
Sources#
Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.
- Benchmark CLI — including the warm-cache warning and the latency definitions
- Parameter sweeps
- Optimization and tuning
- Conserving memory
- GuideLLM
vllm/config/vllm.py—performance_mode