Pidoku

Tuning and Benchmarking

Expert 50 min Difficulty 4/5 Lesson 04 of 05

Prerequisites Chunked Prefill and the Token Budget, Sizing the Cache, Metrics

The idea in one minute#

vLLM has hundreds of flags and perhaps a dozen that matter. Tuning is not trying them at random; it is deciding what you are optimising for, measuring with a load that looks like yours, and changing the one setting that controls the mechanism you found to be the bottleneck. The earlier lessons gave you the mechanisms. This one gives the measurement method — including the single most common benchmarking mistake, which makes an overloaded server look healthy — and a table from symptom to flag.

A picture#

flowchart LR
  G[":i-scale: <b>1. Choose the goal</b><br/><small>latency SLO, throughput, or cost</small>"] --> W[":i-file-text: <b>2. Describe the workload</b><br/><small>prompt and output length distributions,<br/>arrival rate, shared prefixes</small>"]
  W --> B[":i-gauge: <b>3. Measure</b><br/><small>open-loop load at several rates</small>"]
  B --> K[":i-search: <b>4. Find the bottleneck</b><br/><small>metrics: queue, cache, step time, CPU</small>"]
  K --> F[":i-wrench: <b>5. Change one setting</b>"]
  F --> B
  class G,W neutral
  class B compute
  class K queue
  class F warn

How it really works#

Three goals that pull apart#

GoalYou care aboutTypical workloadDirection of the main knobs
InteractivityTime to first token, even streamingChat, coding assistantsSmaller token budget, fine-grained CUDA graphs, headroom in the cache
ThroughputTokens per second per GPUBatch jobs, offline evaluation, agentsLarger token budget and max_num_seqs, full cache
Cost under an SLORequests per GPU that still meet a latency targetMost production servicesTune for throughput until the SLO breaks, then back off

vLLM exposes the first two directly:

Shell
vllm serve <model> --performance-mode interactivity   # or: throughput, balanced (default)

throughput doubles max_num_batched_tokens and max_num_seqs and prefers larger CUDA graphs and throughput-oriented kernels. interactivity captures a CUDA graph for every batch size from 1 to 32 and prefers latency-oriented kernels (CUDA Graphs and Compilation). Start from the mode that matches your goal and adjust from there.

The flags that matter#

FlagMechanismRaise it to…Cost of raising it
--max-num-batched-tokensTokens per stepSpeed up prefill, raise throughputLonger steps: worse inter-token latency; less KV cache
--max-num-seqsRunning requests per engineAdmit more concurrencyBigger batches, more preemption risk, more graph memory
--long-prefill-token-thresholdPer-request share of a step(lower it) to protect short promptsLong prompts take more steps
--gpu-memory-utilizationSize of the block poolAdd cache capacityLess safety margin
--max-model-lenLongest requestAccept longer requestsLower worst-case concurrency; may not start
--kv-cache-dtype fp8Bytes per cached token(set it) to double cache capacitySmall quality risk
--watermarkAdmission headroomReduce preemptionSlightly lower concurrency
--max-num-queued-reqs, --max-num-queued-tokensAdmission control(set them) to bound latency under overloadSome requests get 503
--stream-intervalTokens per streamed eventCut per-token CPU overheadChunkier streaming
--api-server-countFrontend processesRemove a tokenising or media bottleneckCPU cores
-O, --performance-modeCompilation and graph ladder—Startup time and memory
--tensor-parallel-size, --data-parallel-sizeGPU layout—See Parallelism

Measuring: the built-in tools#

Shell
vllm bench serve \
  --model <model> --endpoint /v1/chat/completions \
  --dataset-name random --random-input-len 1024 --random-output-len 256 \
  --num-prompts 2000 --request-rate 20

It reports, measured at the client:

Request throughput (req/s):              1.73
Output token throughput (tok/s):         382.89
Mean TTFT (ms):                          71.54     P99 TTFT (ms):   79.49
Mean TPOT (ms):                          7.91      P99 TPOT (ms):   8.03
Mean ITL (ms):                           7.74      P99 ITL (ms):    8.39
CommandUse
vllm bench serveLoad a running server. Datasets: random, ShareGPT, long-document, multimodal, a prefix-caching workload, timed trace replay.
vllm bench throughputOffline: the LLM class with no server. An upper bound, not a prediction of serving performance.
vllm bench latencyOffline: one request at a time.
vllm bench sweep serveStart the server and run bench serve across a grid of server and client parameters, resetting caches between runs.

For production-grade load testing the documentation points to GuideLLM, a separate project under the same organisation.

Three load-shaping options decide whether your results mean anything:

  • --request-rate R: an open-loop load. Requests arrive at R per second regardless of how the server is doing. inf sends everything at once.
  • --max-concurrency C: a closed-loop limit. At most C requests in flight; a new one starts only when one finishes.
  • --burstiness: 1.0 gives Poisson arrivals; lower values are burstier.

The mistake: closed-loop benchmarks hide overload#

With only --max-concurrency 64, sixty-four clients each send a request, wait for the answer, and send the next. If the server slows down, the clients slow down with it. The arrival rate adapts to the server. No queue can ever build beyond 64, so latency looks stable at any “load” — because you are not controlling load at all.

Real users do not wait for each other. They arrive at their own rate. If that rate exceeds what the server can do, the queue grows without bound and time to first token climbs every second. A closed-loop test cannot show that. It is called coordinated omission, and it is the reason teams ship servers that “benchmarked fine at 64 concurrency” and fall over in production.

The fix: drive load with --request-rate at several values and plot latency against rate. The curve is flat, then bends sharply. The rate just before the bend is the server’s capacity for your SLO. --max-concurrency may be added as a safety cap, but the rate must be the independent variable.

Other ways to fool yourself#

  • Warm caches. The documentation warns: “Repeating vllm bench serve against the same server can reuse prompts left in the prefix cache and inflate throughput.” Vary --seed, restart, or use the sweep tool, which resets caches.
  • The wrong length distribution. Fixed 1,024-in, 256-out traffic has no long prompts to disturb decoders and no variance to cause preemption. Use lengths sampled from your logs: the request-shape histograms in Metrics give them.
  • No shared prefixes, or too many. A random dataset has a 0% prefix-cache hit rate; a dataset that repeats one prompt has 100%. Neither is your workload.
  • Offline numbers for an online service. The LLM class gets a larger token budget by default and has no HTTP, tokenising-under-load or streaming cost.
  • Measuring the client. A Python load generator on one core saturates before a fast server does. Watch the client’s CPU.
  • Averages. Mean latency hides the tail that users complain about. Report p50, p95, p99.
  • Short runs. The first seconds include queue build-up and cold caches. Run long enough to reach steady state, and check num_requests_waiting is not still rising when you stop.

From symptom to setting#

Symptom (with the metric)MechanismChange
High TTFT, high request_queue_time, running = max_num_seqsSequence limitRaise --max-num-seqs
High TTFT, high queue time, kv_cache_usage_perc ≈ 1MemoryFP8 cache, quantised weights, lower --max-model-len, more GPUs
num_preemptions risingOver-admissionLower --max-num-seqs, set --watermark 0.05
High TTFT, low queue time, long promptsPrefill speedRaise --max-num-batched-tokens; fix prompt order for cache hits
High TTFT, low queue and prefill timeFrontend CPUMore --api-server-count, more cores; check media loading
p99 inter-token latency ≫ p50Prefill chunks in decode stepsLower --max-num-batched-tokens; --long-prefill-token-threshold; disaggregate
Short prompts wait behind long onesBudget monopolised--long-prefill-token-threshold 2048
All inter-token latency high, GPU not fullCPU-bound engine coreDedicated physical core; --stream-interval; fewer structured-output requests per engine
Slow at 1 to 4 users, fine at 64Padding and launch overhead--performance-mode interactivity
Latency grows without bound under loadNo admission control--max-num-queued-reqs or --max-num-queued-tokens
Throughput flat while adding concurrencyCompute-saturatedAdd replicas; quantise; this GPU is full
Startup takes minutesCompile and capturePersist ~/.cache/vllm; -O1; vllm preload

A tuning session, in order#

  1. Fit. Make the model start with the max_model_len your traffic needs. Read the capacity line; check capacity ÷ p95 request tokens comfortably exceeds max_num_seqs.
  2. Baseline. Default flags, your length distribution, open-loop load at a low rate. Record p50/p95/p99 TTFT and inter-token latency.
  3. Find the knee. Raise the rate in steps until p95 TTFT or p99 inter-token latency breaks your SLO. Note what the metrics say is binding there.
  4. Move the binding constraint. One flag, from the table. Re-run step 3.
  5. Stop when the constraint is the GPU’s compute, or when the remaining changes trade one SLO for another.
  6. Protect it. Set admission limits a little above the knee so that overload produces fast 503s instead of slow timeouts, and configure the autoscaler on num_requests_waiting or queue time rather than GPU utilisation, which is high whenever the server is busy at all.

Code#

Why a closed-loop test cannot find the knee. The same server is measured two ways: a fixed number of clients, and a fixed arrival rate.

Go
package main

import "fmt"

// The server: it can work on up to maxRunning requests at once, each needing
// serviceSeconds of attention when running alone; with n running, each
// progresses at min(1, capacity/n) of full speed.
const (
	maxRunning     = 64
	capacity       = 32.0 // requests that can run at full speed simultaneously
	serviceSeconds = 2.0
	dt             = 0.01
	duration       = 300.0
)

type req struct {
	arrived, started float64
	left             float64
}

func simulate(arrivalRate float64, closedClients int) (throughput, meanLatency, worstQueueWait float64) {
	var waiting, running []*req
	done, totalLatency, carry := 0, 0.0, 0.0

	if closedClients > 0 {
		for i := 0; i < closedClients; i++ {
			waiting = append(waiting, &req{left: serviceSeconds})
		}
	}
	for t := 0.0; t < duration; t += dt {
		if closedClients == 0 { // open loop: arrivals do not care how the server is doing
			carry += arrivalRate * dt
			for ; carry >= 1; carry-- {
				waiting = append(waiting, &req{arrived: t, left: serviceSeconds})
			}
		}
		for len(waiting) > 0 && len(running) < maxRunning {
			r := waiting[0]
			waiting = waiting[1:]
			r.started = t
			worstQueueWait = max(worstQueueWait, t-r.arrived)
			running = append(running, r)
		}
		speed := min(1.0, capacity/float64(max(len(running), 1)))
		keep := running[:0]
		for _, r := range running {
			if r.left -= speed * dt; r.left > 0 {
				keep = append(keep, r)
				continue
			}
			done++
			totalLatency += t - r.arrived
			if closedClients > 0 { // closed loop: the client sends its next request only now
				waiting = append(waiting, &req{arrived: t, left: serviceSeconds})
			}
		}
		running = keep
	}
	return float64(done) / duration, totalLatency / float64(max(done, 1)), worstQueueWait
}

func main() {
	fmt.Printf("server capacity: %.0f requests/s at full speed\n\n", capacity/serviceSeconds)

	fmt.Println("closed loop (--max-concurrency only)")
	fmt.Printf("  %8s %12s %14s %18s\n", "clients", "req/s", "mean latency", "worst queue wait")
	for _, c := range []int{8, 32, 64, 128, 256} {
		tp, lat, q := simulate(0, c)
		fmt.Printf("  %8d %12.1f %12.2f s %16.2f s\n", c, tp, lat, q)
	}

	fmt.Println("\nopen loop (--request-rate)")
	fmt.Printf("  %8s %12s %14s %18s\n", "rate", "req/s", "mean latency", "worst queue wait")
	for _, r := range []float64{4, 8, 12, 15, 17, 20, 24} {
		tp, lat, q := simulate(r, 0)
		fmt.Printf("  %8.0f %12.1f %12.2f s %16.2f s\n", r, tp, lat, q)
	}
}

In the closed-loop table, latency grows with the number of clients but every row is a steady, finite number, and throughput sits at capacity: 256 clients look like a slower but stable server. In the open-loop table, everything is fine up to the capacity and then the worst queue wait grows with the length of the test, because the queue never stops growing. Only the second table tells you where the cliff is.

Remember this#

  • Decide the goal first: interactivity, throughput, or cost under an SLO. --performance-mode sets a starting point.
  • About a dozen flags matter, and each controls one mechanism from earlier lessons.
  • Benchmark with open-loop load (--request-rate) at several rates; closed-loop tests hide overload.
  • Use your real length distribution and prefix-sharing pattern; reset or vary to avoid warm-cache inflation.
  • Report tail percentiles, not means.
  • Use the metrics to name the binding constraint, change one flag, and re-measure.
  • Set admission limits just above the knee, and autoscale on queue signals, not GPU utilisation.

Try it#

  1. In the program, raise maxRunning to 128 with the same capacity. What happens to open-loop latency just below the knee? This is the cost of a max_num_seqs that is too high for the hardware.
  2. Add admission control to the open-loop case: reject an arrival when more than 40 are waiting. Report the rejection rate and the worst queue wait at a rate of 20.
  3. Run vllm bench serve against a real server at five request rates and plot p95 TTFT. Mark your SLO and read off the capacity.

Check yourself#

  1. Why does a benchmark using only --max-concurrency fail to reveal an overloaded server?
  2. p99 inter-token latency is six times the median while GPU utilisation is modest. Which mechanism is the likely cause and which two flags address it?
  3. Why is GPU utilisation a poor autoscaling signal for an inference server?

Sources#

Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom