The idea in one minute#
vllm serve <model> downloads a model, loads it onto the GPU, measures how much memory is
left, turns that leftover memory into KV cache blocks, compiles and warms up the model, and
only then opens port 8000. The log it prints while doing this is not noise: it states the
exact numbers that the scheduler and memory manager will use for the life of the process.
Learn to read six lines of it and you can predict how the server will behave before sending
a single request.
A picture#
flowchart LR A[":i-terminal: <b>vllm serve</b><br/><small>parse flags, build config</small>"] --> B[":huggingface: <b>Download</b><br/><small>config, tokenizer, weights</small>"] B --> C[":pytorch: <b>Load weights</b><br/><small>onto the GPU</small>"] C --> D[":i-gauge: <b>Profile memory</b><br/><small>one dummy forward pass</small>"] D --> E[":i-layers: <b>Create KV cache</b><br/><small>leftover memory / block size</small>"] E --> F[":nvidia: <b>Compile + CUDA graphs</b><br/><small>warm up every batch shape</small>"] F --> G[":i-globe: <b>Listen on :8000</b><br/><small>OpenAI-compatible API</small>"] class A neutral class B io class C,E memory class D queue class F compute class G io
How it really works#
Install#
vLLM is a Python package. The project recommends uv, which picks the PyTorch build that
matches your GPU driver:
uv venv --python 3.12 --seed
source .venv/bin/activate
uv pip install vllm --torch-backend=auto| Your machine | What to install |
|---|---|
| Linux + NVIDIA GPU | The command above. Supported Python: 3.10 to 3.13 |
| Linux + AMD GPU | uv pip install vllm --extra-index-url https://wheels.vllm.ai/rocm/ (Python 3.12, ROCm 7.0) |
| Linux, CPU only | The +cpu wheel attached to each GitHub release, with --torch-backend cpu |
| Apple Silicon | A separate project, vllm-metal, which uses MLX instead of PyTorch |
| Docker | vllm/vllm-openai:latest; arguments after the image name go to vllm serve |
With Docker:
docker run --rm --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 --ipc=host \
vllm/vllm-openai:latest \
Qwen/Qwen2.5-1.5B-Instruct--ipc=host matters: vLLM’s processes talk to each other through shared memory and local
sockets, and Docker’s default 64 MB shared-memory segment is too small.
Start it#
vllm serve Qwen/Qwen2.5-1.5B-InstructThat is the whole command. One process serves one model. The address defaults to
http://localhost:8000; --host, --port and --uds (a Unix socket) change it.
--api-key or the VLLM_API_KEY environment variable turns on a bearer-token check, and
several keys may be given so that one can be rotated out.
The vllm command has more than serve:
| Command | Purpose |
|---|---|
vllm serve | The HTTP server. This course is mostly about what it starts. |
vllm chat, vllm complete | A terminal client for a running server |
vllm bench latency / throughput / serve | Built-in load generators (see Tuning and Benchmarking) |
vllm run-batch | Process a file of requests offline, in the OpenAI batch format |
vllm launch render | A GPU-less server that only tokenises and applies chat templates |
vllm preload | Load a model into a state that later serve processes can start from quickly |
vllm collect-env | Print versions for a bug report |
vllm serve --help=all prints every flag; --help=max-num-seqs prints one;
--help=ModelConfig prints one group.
Read the startup log#
The lines below are the ones that carry numbers. The wording is taken from the source; the values are an example for a small model on a 24 GB GPU.
non-default args: {'model': 'Qwen/Qwen2.5-1.5B-Instruct'}
Initializing a V1 LLM engine (v0.30.0) with config: model='Qwen/Qwen2.5-1.5B-Instruct', ...
Chunked prefill is enabled with max_num_batched_tokens=2048.
Using max model len 32768
Loading weights took 1.42 seconds
Model loading took 2.9 GiB memory and 2.31 seconds
Available KV cache memory: 18.4 GiB
GPU KV cache size: 688,304 tokens, Maximum concurrency for 32,768 tokens per request: 21.01x
Graph capturing finished in 9 secs, took 0.31 GiB
init engine (profile, create kv cache, warmup model) took 14.87 s
Starting vLLM server on http://0.0.0.0:8000What each one tells you:
| Line | What it means | What to do with it |
|---|---|---|
non-default args | Every flag that differs from the default. Secrets are redacted. | Paste this into bug reports. It is the reproducible configuration. |
Initializing a V1 LLM engine | The engine process has started and prints the whole resolved configuration. | Search it for the value of any flag you are unsure about. |
Chunked prefill is enabled with max_num_batched_tokens=… | The token budget per step: the most tokens the scheduler hands the GPU in one forward pass. | Covered in The Scheduler. The default depends on your GPU. |
Using max model len | The longest prompt plus answer one request may have. Read from the model’s config unless you pass --max-model-len. | Lowering it raises the “maximum concurrency” number below. |
Model loading took … GiB | Memory the weights occupy. | Subtract from the GPU to estimate what is left. |
Available KV cache memory | What remains inside the budget after weights and peak activation memory. | This number divided by bytes-per-block is the size of the block pool. |
GPU KV cache size: N tokens | The block pool expressed in tokens. | The total of all tokens, across all requests, that can be in flight at once. |
Maximum concurrency for M tokens per request: X | N / M: how many requests could run if every one used the full context. | A worst-case floor, not a limit. Real requests are shorter, so real concurrency is higher. |
Graph capturing finished | CUDA graphs were recorded for a set of batch sizes. | Startup time goes here. See CUDA Graphs and Compilation. |
init engine … took | Total time for profiling, cache creation and warm-up. | The floor on your cold-start time, on top of download and load. |
important
The budget is --gpu-memory-utilization, which defaults to 0.92 of the GPU’s total
memory. vLLM takes that share at startup and keeps it. A vLLM process showing 22 GB used
on a 24 GB card while idle is working as designed: the blocks are pre-allocated so that no
allocation ever happens while serving.
Talk to it#
The API is OpenAI’s. Anything that can call OpenAI can call vLLM by changing the base URL.
curl http://localhost:8000/v1/models
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"messages": [{"role": "user", "content": "Name three uses of a GPU."}],
"max_tokens": 64
}'The endpoints a generation server exposes:
| Path | What it is |
|---|---|
/v1/chat/completions | Messages in, message out. Applies the model’s chat template. |
/v1/completions | Raw text in, text out. No template. |
/v1/responses | OpenAI’s Responses API |
/v1/messages | Anthropic’s Messages API, served by the same engine |
/v1/models | The model name and any loaded LoRA adapters |
/tokenize, /detokenize | Run only the tokenizer |
/health, /ping | Liveness; returns 200 once the engine is up |
/metrics | Prometheus metrics (see Metrics) |
/version | The vLLM version |
Embedding, reranking, classification and speech models get their own endpoints
(/v1/embeddings, /rerank, /classify, /v1/audio/transcriptions) when you serve a model
of that kind.
note
By default the server applies the sampling defaults in the model repository’s
generation_config.json. If answers differ from what you expect at “default” settings,
that file is usually why. Pass --generation-config vllm to ignore it.
Two ways to use the engine#
The HTTP server is one of two front doors to the same engine:
flowchart LR S[":i-globe: <b>vllm serve</b><br/><small>online: many clients, streaming</small>"] --> A[":vllm: <b>AsyncLLM</b>"] L[":python: <b>LLM class</b><br/><small>offline: a list of prompts</small>"] --> B[":vllm: <b>LLMEngine</b>"] A --> E[":vllm: <b>Engine core</b><br/><small>the same scheduler and workers</small>"] B --> E class S,L io class A,B queue class E compute
The LLM class is for batch jobs inside a Python script: give it ten thousand prompts and it
returns ten thousand answers. Note the larger defaults it gets — on an H100-class GPU the
offline token budget is 16,384 per step against 8,192 for the server — because a batch job
cares only about throughput.
Code#
A Go client that streams a chat completion. It needs a running server, so start one first.
The response is server-sent events: lines beginning data: , each carrying a JSON chunk
with the next piece of text, ending with data: [DONE].
package main
import (
"bufio"
"bytes"
"encoding/json"
"fmt"
"net/http"
"os"
"strings"
"time"
)
type chunk struct {
Choices []struct {
Delta struct {
Content string `json:"content"`
} `json:"delta"`
FinishReason *string `json:"finish_reason"`
} `json:"choices"`
Usage *struct {
PromptTokens int `json:"prompt_tokens"`
CompletionTokens int `json:"completion_tokens"`
} `json:"usage"`
}
func main() {
base := "http://localhost:8000"
if v := os.Getenv("VLLM_URL"); v != "" {
base = v
}
body, _ := json.Marshal(map[string]any{
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"messages": []map[string]string{{"role": "user", "content": "Explain a KV cache in two sentences."}},
"max_tokens": 128,
"stream": true,
"stream_options": map[string]bool{"include_usage": true},
})
start := time.Now()
resp, err := http.Post(base+"/v1/chat/completions", "application/json", bytes.NewReader(body))
if err != nil {
fmt.Println("is the server running?", err)
return
}
defer resp.Body.Close()
var first time.Duration
tokens := 0
sc := bufio.NewScanner(resp.Body)
for sc.Scan() {
line := strings.TrimPrefix(sc.Text(), "data: ")
if line == "" || line == "[DONE]" {
continue
}
var c chunk
if json.Unmarshal([]byte(line), &c) != nil {
continue
}
if c.Usage != nil {
tokens = c.Usage.CompletionTokens
}
if len(c.Choices) == 0 {
continue
}
if first == 0 && c.Choices[0].Delta.Content != "" {
first = time.Since(start)
}
fmt.Print(c.Choices[0].Delta.Content)
}
total := time.Since(start)
fmt.Printf("\n\ntime to first token: %v\n", first.Round(time.Millisecond))
fmt.Printf("total: %v for %d tokens\n", total.Round(time.Millisecond), tokens)
if tokens > 1 {
fmt.Printf("per token after the first: %v\n",
((total - first) / time.Duration(tokens-1)).Round(time.Microsecond))
}
}The two numbers it prints are the two that matter for an interactive client: time to first token (queueing plus reading the prompt) and time per output token (one decode step). Every later lesson explains a mechanism that moves one or both.
Remember this#
vllm serve <model>is the whole command; one process, one model, port 8000.- Startup is: load weights, profile memory, create the block pool, compile and warm up, listen.
- vLLM takes 92% of the GPU by default and keeps it; that is the block pool, not a leak.
GPU KV cache sizein tokens is the real capacity.Maximum concurrencyis the worst case.non-default argsis the line to paste into a bug report.- The same engine sits behind the HTTP server (
AsyncLLM) and the offlineLLMclass.
Try it#
- Start the server twice, once with
--max-model-len 4096and once with--max-model-len 32768. The KV cache size in tokens is the same both times. Why does the maximum concurrency change by a factor of eight? - Start it with
--gpu-memory-utilization 0.5. CompareAvailable KV cache memorywith the default run. Is it half? Explain the difference using the weights’ memory. - Run the Go client three times with the same prompt. The time to first token usually drops after the first run. Form a guess about why; the answer is in Prefix Caching.
Check yourself#
- Which startup line tells you the scheduler’s token budget per step?
- An idle vLLM process holds most of the GPU’s memory. Is something wrong?
- What is the difference between
/v1/chat/completionsand/v1/completions?
Sources#
Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.
- Quickstart
- vLLM CLI guide
- Online serving: supported APIs
vllm/config/cache.py— thegpu_memory_utilizationdefaultvllm/v1/core/kv_cache_utils.py— the capacity log line