Pidoku
02Your First Server, and How to Read Its Startup Log
On this page

Your First Server, and How to Read Its Startup Log

Foundations 40 min Difficulty 1/5 Lesson 02 of 03

Prerequisites What vLLM Is and Why It Exists

The idea in one minute#

vllm serve <model> downloads a model, loads it onto the GPU, measures how much memory is left, turns that leftover memory into KV cache blocks, compiles and warms up the model, and only then opens port 8000. The log it prints while doing this is not noise: it states the exact numbers that the scheduler and memory manager will use for the life of the process. Learn to read six lines of it and you can predict how the server will behave before sending a single request.

A picture#

flowchart LR
  A[":i-terminal: <b>vllm serve</b><br/><small>parse flags, build config</small>"] --> B[":huggingface: <b>Download</b><br/><small>config, tokenizer, weights</small>"]
  B --> C[":pytorch: <b>Load weights</b><br/><small>onto the GPU</small>"]
  C --> D[":i-gauge: <b>Profile memory</b><br/><small>one dummy forward pass</small>"]
  D --> E[":i-layers: <b>Create KV cache</b><br/><small>leftover memory / block size</small>"]
  E --> F[":nvidia: <b>Compile + CUDA graphs</b><br/><small>warm up every batch shape</small>"]
  F --> G[":i-globe: <b>Listen on :8000</b><br/><small>OpenAI-compatible API</small>"]
  class A neutral
  class B io
  class C,E memory
  class D queue
  class F compute
  class G io

How it really works#

Install#

vLLM is a Python package. The project recommends uv, which picks the PyTorch build that matches your GPU driver:

Shell
uv venv --python 3.12 --seed
source .venv/bin/activate
uv pip install vllm --torch-backend=auto
Your machineWhat to install
Linux + NVIDIA GPUThe command above. Supported Python: 3.10 to 3.13
Linux + AMD GPUuv pip install vllm --extra-index-url https://wheels.vllm.ai/rocm/ (Python 3.12, ROCm 7.0)
Linux, CPU onlyThe +cpu wheel attached to each GitHub release, with --torch-backend cpu
Apple SiliconA separate project, vllm-metal, which uses MLX instead of PyTorch
Dockervllm/vllm-openai:latest; arguments after the image name go to vllm serve

With Docker:

Shell
docker run --rm --gpus all \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    -p 8000:8000 --ipc=host \
    vllm/vllm-openai:latest \
    Qwen/Qwen2.5-1.5B-Instruct

--ipc=host matters: vLLM’s processes talk to each other through shared memory and local sockets, and Docker’s default 64 MB shared-memory segment is too small.

Start it#

Shell
vllm serve Qwen/Qwen2.5-1.5B-Instruct

That is the whole command. One process serves one model. The address defaults to http://localhost:8000; --host, --port and --uds (a Unix socket) change it. --api-key or the VLLM_API_KEY environment variable turns on a bearer-token check, and several keys may be given so that one can be rotated out.

The vllm command has more than serve:

CommandPurpose
vllm serveThe HTTP server. This course is mostly about what it starts.
vllm chat, vllm completeA terminal client for a running server
vllm bench latency / throughput / serveBuilt-in load generators (see Tuning and Benchmarking)
vllm run-batchProcess a file of requests offline, in the OpenAI batch format
vllm launch renderA GPU-less server that only tokenises and applies chat templates
vllm preloadLoad a model into a state that later serve processes can start from quickly
vllm collect-envPrint versions for a bug report

vllm serve --help=all prints every flag; --help=max-num-seqs prints one; --help=ModelConfig prints one group.

Read the startup log#

The lines below are the ones that carry numbers. The wording is taken from the source; the values are an example for a small model on a 24 GB GPU.

non-default args: {'model': 'Qwen/Qwen2.5-1.5B-Instruct'}
Initializing a V1 LLM engine (v0.30.0) with config: model='Qwen/Qwen2.5-1.5B-Instruct', ...
Chunked prefill is enabled with max_num_batched_tokens=2048.
Using max model len 32768
Loading weights took 1.42 seconds
Model loading took 2.9 GiB memory and 2.31 seconds
Available KV cache memory: 18.4 GiB
GPU KV cache size: 688,304 tokens, Maximum concurrency for 32,768 tokens per request: 21.01x
Graph capturing finished in 9 secs, took 0.31 GiB
init engine (profile, create kv cache, warmup model) took 14.87 s
Starting vLLM server on http://0.0.0.0:8000

What each one tells you:

LineWhat it meansWhat to do with it
non-default argsEvery flag that differs from the default. Secrets are redacted.Paste this into bug reports. It is the reproducible configuration.
Initializing a V1 LLM engineThe engine process has started and prints the whole resolved configuration.Search it for the value of any flag you are unsure about.
Chunked prefill is enabled with max_num_batched_tokens=…The token budget per step: the most tokens the scheduler hands the GPU in one forward pass.Covered in The Scheduler. The default depends on your GPU.
Using max model lenThe longest prompt plus answer one request may have. Read from the model’s config unless you pass --max-model-len.Lowering it raises the “maximum concurrency” number below.
Model loading took … GiBMemory the weights occupy.Subtract from the GPU to estimate what is left.
Available KV cache memoryWhat remains inside the budget after weights and peak activation memory.This number divided by bytes-per-block is the size of the block pool.
GPU KV cache size: N tokensThe block pool expressed in tokens.The total of all tokens, across all requests, that can be in flight at once.
Maximum concurrency for M tokens per request: XN / M: how many requests could run if every one used the full context.A worst-case floor, not a limit. Real requests are shorter, so real concurrency is higher.
Graph capturing finishedCUDA graphs were recorded for a set of batch sizes.Startup time goes here. See CUDA Graphs and Compilation.
init engine … tookTotal time for profiling, cache creation and warm-up.The floor on your cold-start time, on top of download and load.

important

The budget is --gpu-memory-utilization, which defaults to 0.92 of the GPU’s total memory. vLLM takes that share at startup and keeps it. A vLLM process showing 22 GB used on a 24 GB card while idle is working as designed: the blocks are pre-allocated so that no allocation ever happens while serving.

Talk to it#

The API is OpenAI’s. Anything that can call OpenAI can call vLLM by changing the base URL.

Shell
curl http://localhost:8000/v1/models

curl http://localhost:8000/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "Qwen/Qwen2.5-1.5B-Instruct",
        "messages": [{"role": "user", "content": "Name three uses of a GPU."}],
        "max_tokens": 64
    }'

The endpoints a generation server exposes:

PathWhat it is
/v1/chat/completionsMessages in, message out. Applies the model’s chat template.
/v1/completionsRaw text in, text out. No template.
/v1/responsesOpenAI’s Responses API
/v1/messagesAnthropic’s Messages API, served by the same engine
/v1/modelsThe model name and any loaded LoRA adapters
/tokenize, /detokenizeRun only the tokenizer
/health, /pingLiveness; returns 200 once the engine is up
/metricsPrometheus metrics (see Metrics)
/versionThe vLLM version

Embedding, reranking, classification and speech models get their own endpoints (/v1/embeddings, /rerank, /classify, /v1/audio/transcriptions) when you serve a model of that kind.

note

By default the server applies the sampling defaults in the model repository’s generation_config.json. If answers differ from what you expect at “default” settings, that file is usually why. Pass --generation-config vllm to ignore it.

Two ways to use the engine#

The HTTP server is one of two front doors to the same engine:

flowchart LR
  S[":i-globe: <b>vllm serve</b><br/><small>online: many clients, streaming</small>"] --> A[":vllm: <b>AsyncLLM</b>"]
  L[":python: <b>LLM class</b><br/><small>offline: a list of prompts</small>"] --> B[":vllm: <b>LLMEngine</b>"]
  A --> E[":vllm: <b>Engine core</b><br/><small>the same scheduler and workers</small>"]
  B --> E
  class S,L io
  class A,B queue
  class E compute

The LLM class is for batch jobs inside a Python script: give it ten thousand prompts and it returns ten thousand answers. Note the larger defaults it gets — on an H100-class GPU the offline token budget is 16,384 per step against 8,192 for the server — because a batch job cares only about throughput.

Code#

A Go client that streams a chat completion. It needs a running server, so start one first. The response is server-sent events: lines beginning data: , each carrying a JSON chunk with the next piece of text, ending with data: [DONE].

Go
package main

import (
	"bufio"
	"bytes"
	"encoding/json"
	"fmt"
	"net/http"
	"os"
	"strings"
	"time"
)

type chunk struct {
	Choices []struct {
		Delta struct {
			Content string `json:"content"`
		} `json:"delta"`
		FinishReason *string `json:"finish_reason"`
	} `json:"choices"`
	Usage *struct {
		PromptTokens     int `json:"prompt_tokens"`
		CompletionTokens int `json:"completion_tokens"`
	} `json:"usage"`
}

func main() {
	base := "http://localhost:8000"
	if v := os.Getenv("VLLM_URL"); v != "" {
		base = v
	}
	body, _ := json.Marshal(map[string]any{
		"model":          "Qwen/Qwen2.5-1.5B-Instruct",
		"messages":       []map[string]string{{"role": "user", "content": "Explain a KV cache in two sentences."}},
		"max_tokens":     128,
		"stream":         true,
		"stream_options": map[string]bool{"include_usage": true},
	})

	start := time.Now()
	resp, err := http.Post(base+"/v1/chat/completions", "application/json", bytes.NewReader(body))
	if err != nil {
		fmt.Println("is the server running?", err)
		return
	}
	defer resp.Body.Close()

	var first time.Duration
	tokens := 0
	sc := bufio.NewScanner(resp.Body)
	for sc.Scan() {
		line := strings.TrimPrefix(sc.Text(), "data: ")
		if line == "" || line == "[DONE]" {
			continue
		}
		var c chunk
		if json.Unmarshal([]byte(line), &c) != nil {
			continue
		}
		if c.Usage != nil {
			tokens = c.Usage.CompletionTokens
		}
		if len(c.Choices) == 0 {
			continue
		}
		if first == 0 && c.Choices[0].Delta.Content != "" {
			first = time.Since(start)
		}
		fmt.Print(c.Choices[0].Delta.Content)
	}
	total := time.Since(start)

	fmt.Printf("\n\ntime to first token: %v\n", first.Round(time.Millisecond))
	fmt.Printf("total:               %v for %d tokens\n", total.Round(time.Millisecond), tokens)
	if tokens > 1 {
		fmt.Printf("per token after the first: %v\n",
			((total - first) / time.Duration(tokens-1)).Round(time.Microsecond))
	}
}

The two numbers it prints are the two that matter for an interactive client: time to first token (queueing plus reading the prompt) and time per output token (one decode step). Every later lesson explains a mechanism that moves one or both.

Remember this#

  • vllm serve <model> is the whole command; one process, one model, port 8000.
  • Startup is: load weights, profile memory, create the block pool, compile and warm up, listen.
  • vLLM takes 92% of the GPU by default and keeps it; that is the block pool, not a leak.
  • GPU KV cache size in tokens is the real capacity. Maximum concurrency is the worst case.
  • non-default args is the line to paste into a bug report.
  • The same engine sits behind the HTTP server (AsyncLLM) and the offline LLM class.

Try it#

  1. Start the server twice, once with --max-model-len 4096 and once with --max-model-len 32768. The KV cache size in tokens is the same both times. Why does the maximum concurrency change by a factor of eight?
  2. Start it with --gpu-memory-utilization 0.5. Compare Available KV cache memory with the default run. Is it half? Explain the difference using the weights’ memory.
  3. Run the Go client three times with the same prompt. The time to first token usually drops after the first run. Form a guess about why; the answer is in Prefix Caching.

Check yourself#

  1. Which startup line tells you the scheduler’s token budget per step?
  2. An idle vLLM process holds most of the GPU’s memory. Is something wrong?
  3. What is the difference between /v1/chat/completions and /v1/completions?

Sources#

Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom