Pidoku

What vLLM Is and Why It Exists

Foundations 30 min Difficulty 1/5 Lesson 01 of 03

Prerequisites You can run a program from a terminal. Nothing else.

The idea in one minute#

A language model produces text one token at a time. Every request therefore needs the GPU thousands of times in a row, and every request must keep a growing pile of intermediate results in GPU memory while it does so. An inference engine is the program that decides, many times per second, which requests get the GPU next and where each request’s memory lives. vLLM is the most widely used open-source inference engine. It is a Python program with C++, CUDA and (recently) Rust inside it; it loads a model from Hugging Face, exposes the same HTTP API as OpenAI, and squeezes far more requests through one GPU than a naive loop can.

It began with one idea — treat the GPU memory that requests use the way an operating system treats RAM, in small fixed-size pages — and everything else in it grew around that idea.

A picture#

flowchart LR
  U[":i-users: <b>Many clients</b><br/><small>chat apps, agents, batch jobs</small>"] -->|"HTTP, OpenAI format"| API[":vllm: <b>API server</b><br/><small>tokenise, validate, stream</small>"]
  API -->|"token IDs"| ENG[":vllm: <b>Engine core</b><br/><small>scheduler + KV cache manager</small>"]
  ENG -->|"this step's batch"| W[":pytorch: <b>GPU worker</b><br/><small>runs the model</small>"]
  W --> GPU[":nvidia: <b>GPU</b><br/><small>weights + KV cache blocks</small>"]
  W -->|"one new token per request"| ENG
  ENG -->|"new tokens"| API
  API -->|"text, streamed"| U
  HF[(":huggingface: <b>Model files</b>")] -.->|"loaded once"| W
  class U neutral
  class API io
  class ENG queue
  class W compute
  class GPU memory
  class HF memory

Three boxes do all the work, and this course takes them apart one at a time: the API server (text in, text out), the engine core (the decisions) and the worker (the arithmetic).

How it really works#

The problem a naive server has#

Suppose you write the obvious server: one request arrives, you run the model until the answer is finished, then you take the next request. Three things go wrong.

  1. The GPU is mostly idle. A GPU is fastest when it processes many sequences in one call. One request at a time uses a small fraction of it.
  2. Batching by hand is rigid. If you wait to collect eight requests and run them together, the short answers finish early and their slots sit empty until the longest one is done.
  3. Memory is wasted. While generating, each request stores a KV cache: saved attention results for every token so far, so that earlier tokens are never recomputed (see KV Cache). You do not know in advance how long an answer will be, so the obvious approach reserves room for the longest possible answer for every request. Most of that reservation is never used.

Point 3 was the one nobody had solved cleanly in 2023. A KV cache for one request can be hundreds of megabytes, and reserving the maximum for everyone meant a 24 GB GPU could hold only a handful of requests even when all of them were short.

The idea vLLM started from: pages#

Operating systems met the same problem decades ago. A process does not get one contiguous slab of RAM sized for its worst case. It gets pages — small, fixed-size pieces handed out only when used — and a page table that records which pages belong to it.

vLLM does this for the KV cache:

Operating systemvLLM
Page (4 KiB)Block: room for the KV of 16 tokens by default
Physical memoryOne large tensor per layer, cut into blocks, allocated at startup
Page tableBlock table: for each request, the list of block numbers it owns
Process grows its heapRequest generates more tokens, takes one more block
Shared read-only pagesTwo requests with the same prompt prefix point at the same blocks
Free listA queue of unused blocks

The attention code was rewritten so it can read a request’s KV cache from blocks scattered anywhere in the tensor, following the block table. That kernel is called PagedAttention, and it is the subject of the paper that introduced vLLM (Kwon et al., SOSP 2023).

With pages, a request holds only the memory for tokens it has actually produced, rounded up to one block. Waste per request is at most 15 tokens’ worth instead of thousands.

What grew around it#

Paging made the second idea practical: continuous batching. Because memory is handed out in small pieces, a new request can join the batch the moment another one finishes, without waiting for the whole batch to drain. The engine re-decides the batch on every step.

Since 2023 the project has added, layer by layer:

LayerWhat it doesWhere in this course
SchedulerChooses which requests run each step under a token budgetThe Scheduler
Block pool and prefix cacheHands out blocks, reuses blocks for repeated prompt prefixesKV Cache
Model runnerTurns a schedule into tensors and calls the modelModel Execution
Attention backendsFlashAttention, FlashInfer and others, chosen per GPU and modelModel Execution
FeaturesSpeculative decoding, LoRA, multimodal, structured output, tool callingFeatures
ParallelismOne model across many GPUs, or many copies across many GPUsScaling and Operating

vLLM by the numbers#

These were read from the project on 5 October 2026. They change every fortnight; the ideas in this course do not.

FactValue
Latest releasev0.30.0, 22 September 2026
Release rhythmA minor version roughly every two weeks (v0.25.0 on 11 July, v0.30.0 on 22 September)
LicenceApache 2.0
GitHub stars / forksabout 93,000 / 23,000
Contributorsmore than 2,000, according to the project README
Python sourcejust over one million lines under vllm/
Supported Python3.10 to 3.13 for the released wheels
PyTorch pinned by the CUDA build2.13.0
Model architectures“200+” on Hugging Face, by the project’s count
OriginSky Computing Lab, UC Berkeley; first public release June 2023

What vLLM is not#

  • It is not a model. It runs models that others trained.
  • It is not a training framework. It only does the forward pass.
  • It is not a gateway. It serves one model (plus adapters) per process; routing between models, rate limits per customer and billing sit in front of it (see Model Access and Gateways).
  • It is not the only engine. SGLang, TGI, TensorRT-LLM and llama.cpp solve the same problem with different trade-offs (see Choosing a Stack).

Why read its source at all#

Most people use vLLM as a black box with forty flags. That works until latency doubles at 3 p.m. and the only tool you have is guessing. Every flag is a constant in a loop you can read. Once you have seen that loop — it is about a thousand lines — the flags stop being folklore:

  • --max-num-batched-tokens is the number the scheduler subtracts from on each step.
  • --gpu-memory-utilization decides how many blocks the pool has.
  • “Preemption” in the logs is one request’s blocks being taken away to feed another.

This course quotes the real code for each of these. Everything was read from the main branch at commit 0c16eee (5 October 2026) and cross-checked against the v0.30.0 tag; where the two differ, the lesson says so.

Code#

Before any vLLM code, measure the problem it solves. This program serves the same requests two ways: reserving the maximum length for each, and handing out 16-token blocks on demand.

Go
package main

import (
	"fmt"
	"math/rand"
)

const (
	maxTokens   = 4096 // the longest sequence the server accepts
	blockSize   = 16   // tokens per block
	gpuCapacity = 1 << 20
)

func main() {
	rng := rand.New(rand.NewSource(7))

	// Real traffic is skewed: most requests are short, a few are long.
	lengths := make([]int, 0, 4000)
	for i := 0; i < 4000; i++ {
		n := 100 + int(rng.ExpFloat64()*400)
		if n > maxTokens {
			n = maxTokens
		}
		lengths = append(lengths, n)
	}

	// Strategy 1: reserve the maximum for every request.
	reserved := gpuCapacity / maxTokens

	// Strategy 2: give each request only the blocks it uses.
	used, paged := 0, 0
	for _, n := range lengths {
		blocks := (n + blockSize - 1) / blockSize
		if used+blocks*blockSize > gpuCapacity {
			break
		}
		used += blocks * blockSize
		paged++
	}

	actual := 0
	for _, n := range lengths[:paged] {
		actual += n
	}

	fmt.Printf("capacity:                 %d tokens of KV cache\n", gpuCapacity)
	fmt.Printf("reserve-the-maximum:      %d requests fit\n", reserved)
	fmt.Printf("paged, 16-token blocks:   %d requests fit\n", paged)
	fmt.Printf("improvement:              %.1fx\n", float64(paged)/float64(reserved))
	fmt.Printf("memory wasted by paging:  %.2f%% (rounding up to a block)\n",
		100*float64(used-actual)/float64(used))
}

Run it. With this traffic the paged server holds roughly eight times as many requests in the same memory, and wastes about one and a half percent on rounding. That gap is the reason vLLM exists.

Remember this#

  • An inference engine decides who gets the GPU each step and where each request’s memory lives.
  • vLLM’s founding idea is paging: the KV cache is cut into fixed-size blocks handed out on demand.
  • Paging removes the need to reserve the maximum length, which is what makes continuous batching practical.
  • vLLM is three cooperating parts: API server, engine core, GPU worker.
  • It serves one model per process and speaks the OpenAI HTTP API.
  • The flags are constants in a readable loop; the rest of this course reads that loop.

Try it#

  1. Change blockSize to 1, 64 and 512 and rerun. What happens to the waste, and why would a block size of 1 be a bad idea even though it wastes nothing? (Think about the size of the block table.)
  2. Change the traffic so every request is exactly 4,096 tokens. What is the improvement now? What does that tell you about when paging helps?
  3. Open the vLLM releases page and find the date of the newest release. How many releases have there been since v0.30.0?

Check yourself#

  1. Why can a server not know how much KV cache memory a request needs when it arrives?
  2. What is a block table, and what is its operating-system equivalent?
  3. Name the three parts of a running vLLM server and what each one is responsible for.

Sources#

Checked on 5 October 2026.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom