Short definitions of the terms this course uses. The lesson in brackets is where each is explained.
| Term | Meaning |
|---|---|
| Admission control | Refusing new requests with HTTP 503 when queue limits are reached. (The Frontend) |
| All-reduce | A collective operation that sums partial results across GPUs; two per layer under tensor parallelism. (Parallelism) |
| APC | Automatic prefix caching. (Prefix Caching) |
| API server | The process that handles HTTP, templates, tokenising and detokenising. (Processes and Wires) |
| Argmax-invariant | A logits transformation that cannot change which token scores highest. (Sampling and Structured Output) |
| Async scheduling | Scheduling step N+1 while step N is still on the GPU. (Async Scheduling) |
AsyncLLM | The engine as seen from an asyncio program; what the server uses. (AsyncLLM and the Way Back) |
| Attention backend | A kernel implementation that can read a paged KV cache, chosen at startup. (Attention Backends) |
| Batch queue | Batches scheduled but not yet finished; size 2 with async scheduling. (The Engine Core Loop) |
| Block | Fixed-size unit of KV cache: 16 tokens by default. (Blocks and the Pool) |
| Block hash | A hash of a block’s tokens and its parent’s hash; identifies a whole prefix. (Prefix Caching) |
| Block pool | All KV blocks, the free queue and the hash-to-block map. (Blocks and the Pool) |
| Block table | A request’s ordered list of physical block IDs. (Blocks and the Pool) |
| Busy loop | The engine core’s main loop; never sleeps while requests exist. (The Engine Core Loop) |
cache_salt | A per-request value hashed into the first block to isolate tenants’ caches. (Prefix Caching) |
| Cascade attention | Computing attention over a prefix shared by all requests once. (Attention Backends) |
| Chat template | A Jinja2 program that renders messages and tools into the model’s prompt format. (The Frontend) |
| Chunked prefill | Reading a long prompt over several steps, bounded by the token budget. (Chunked Prefill and the Token Budget) |
| Closed-loop load | A benchmark where a new request is sent only when one finishes; hides overload. (Tuning and Benchmarking) |
| Collective RPC | Calling a method on every worker and gathering results. (Executor, Worker, Model Runner) |
| Constrained decoding | Masking tokens a grammar forbids before sampling. (Sampling and Structured Output) |
| Continuous batching | Re-deciding the batch every step so requests join and leave freely. (What vLLM Is and Why It Exists) |
| Copy-on-write | Copying a shared, partly matching block into a private one before writing. (Prefix Caching) |
| CUDA graph | A recording of GPU operations for one batch shape, replayed in one launch. (CUDA Graphs and Compilation) |
| Data parallelism (DP) | Several complete engines behind one endpoint. (Parallelism) |
| Deferred free | Holding a finished request’s blocks until in-flight steps that may write them complete. (Async Scheduling) |
| Detokenizer, incremental | Converts token IDs to text step by step, emitting only complete characters. (AsyncLLM and the Way Back) |
| Disaggregation (P/D) | Running prefill and decode on different instances. (Disaggregation and KV Connectors) |
| DP coordinator | Process that balances load and aligns steps across data-parallel ranks. (Parallelism) |
| Encoder cache | Holds a vision or audio encoder’s output until the request has consumed it. (Multimodal Inputs) |
| Engine core | The process that runs the scheduler and KV cache manager. (Processes and Wires) |
EngineCoreRequest / EngineCoreOutput | The msgpack messages between API server and engine core. (Processes and Wires) |
| Eviction | A cached block losing its hash when handed to a new owner; lazy in vLLM. (Blocks and the Pool) |
| Executor | Delivers one call to every worker of an engine. (Executor, Worker, Model Runner) |
| Expert parallelism (EP) | Distributing a mixture-of-experts model’s experts across ranks. (Parallelism) |
| Flat batch | All scheduled tokens of all requests in one row, with index tensors instead of padding. (Building the Batch) |
| Free queue | Doubly linked list of unowned blocks; its order is the eviction policy. (Blocks and the Pool) |
| Gap | Tokens a request has minus tokens computed; what the scheduler closes. (The Map) |
| Grammar bitmask | One bit per vocabulary token saying whether it is legal next. (Sampling and Structured Output) |
| Head-of-line blocking | A waiting request that does not fit stops admission behind it. (One Scheduling Step) |
| Inter-token latency (ITL) | Time between consecutive streamed outputs. (Metrics) |
| KV cache group | Layers with the same cache needs, sharing one block table per request. (More Than One Kind of Cache) |
| KV cache spec | A layer’s declaration of what its cache stores and how large a page is. (Sizing the Cache) |
| KV connector | Plug-in that moves cached blocks to other tiers or machines. (Disaggregation and KV Connectors) |
| KV events | Published notices of blocks stored and removed, for cache-aware routers. (Disaggregation and KV Connectors) |
logits_indices | Rows of the output that become next-token distributions: each request’s last. (Building the Batch) |
| Lookahead | Extra cache slots reserved for a speculative proposer. (Allocating Slots) |
| LoRA | A low-rank correction to a base model, applied per request at run time. (LoRA Adapters) |
max_num_batched_tokens | The token budget: most tokens processed in one step. (Chunked Prefill and the Token Budget) |
max_num_seqs | Most requests running at once, per engine. (One Scheduling Step) |
| Model runner | Builds the batch, runs the model and samples, inside a worker. (Executor, Worker, Model Runner) |
| MRV2 | Model Runner V2: the newer runner with fixed rows and GPU-side input preparation. (Executor, Worker, Model Runner) |
| Null block | Block 0; a placeholder never allocated or freed. (Blocks and the Pool) |
num_computed_tokens | Positions whose KV is computed or scheduled to be. (One Scheduling Step) |
| Offloading | Copying cached blocks to CPU memory or disk to extend the prefix cache. (Disaggregation and KV Connectors) |
| Open-loop load | A benchmark with arrivals at a fixed rate regardless of server state. (Tuning and Benchmarking) |
| Optimisation level | -O0 to -O3: presets trading startup time for speed. (CUDA Graphs and Compilation) |
| Output placeholder | A scheduled position whose token ID is not known yet. (Async Scheduling) |
| Page size | Bytes one block occupies for one layer or group. (Sizing the Cache) |
| PagedAttention | Attention over a KV cache stored in scattered fixed-size blocks. (What vLLM Is and Why It Exists) |
| Persistent batch | Per-request state kept between steps and updated by differences. (Executor, Worker, Model Runner) |
| Piecewise compilation | Splitting the model’s graph at attention calls so the rest can be recorded. (CUDA Graphs and Compilation) |
| Pipeline parallelism (PP) | Splitting the layer stack into stages on different GPUs. (Parallelism) |
| Placeholder token | A token standing in for one vector of an image or audio clip. (Multimodal Inputs) |
| Plugin | Code discovered through Python entry points and loaded in every vLLM process. (Plugins and Extension Points) |
| Preemption | Taking all of a running request’s blocks; it recomputes when re-admitted. (Priorities, Preemption and Queues) |
| Prefix caching | Reusing KV blocks for a prompt prefix seen before. (Prefix Caching) |
prefix_match_unit | Hashing granularity, which may be finer than the physical block. (Prefix Caching) |
query_start_loc | Offsets marking each request’s rows in the flat batch. (Building the Batch) |
| Reasoning parser | Splits a model’s thinking from its answer into a separate field. (Tool Calling and Reasoning) |
| Reference count | Number of requests whose block table contains a block. (Blocks and the Pool) |
| Renderer | Applies the chat template and tokenises. (The Frontend) |
| Request status | WAITING, RUNNING, PREEMPTED, several waiting-for states, and the finished states. (One Scheduling Step) |
| Ring buffer, shared-memory | The lock-free broadcast channel from engine core to workers. (Processes and Wires) |
SchedulerOutput | One step’s plan: which requests, how many tokens, which blocks. (One Scheduling Step) |
| Sleep mode | Releasing GPU memory while keeping the process alive. (The Engine Core Loop) |
| Slot mapping | For each new token, the physical cache slot its key and value are written to. (Building the Batch) |
| Speculative decoding | Proposing several tokens cheaply and verifying them in one pass. (Speculative Decoding) |
| Stale output | A result from a step scheduled before a request was preempted. (Async Scheduling) |
| Step | One schedule-execute-update cycle; at most one token per request without speculation. (The Engine Core Loop) |
| Stop string | Text that ends generation; checked in the API server, not the engine. (AsyncLLM and the Way Back) |
| Structured output | Output guaranteed to match a JSON schema, regex or grammar. (Sampling and Structured Output) |
| Tensor parallelism (TP) | Splitting each layer’s weights across GPUs. (Parallelism) |
| Time per output token (TPOT) | Per request: (end-to-end latency − TTFT) ÷ (output tokens − 1). (Metrics) |
| Time to first token (TTFT) | Arrival to first streamed token. (Metrics) |
| Token budget | See max_num_batched_tokens. |
| Tool parser | Extracts tool calls from a model’s free-text output. (Tool Calling and Reasoning) |
| Touch | A cache hit on a block: reference count up, removed from the free queue. (Blocks and the Pool) |
| Uniform batch | Every request contributing the same number of tokens; eligible for a full CUDA graph. (CUDA Graphs and Compilation) |
| Utility call | A method invoked on the engine core over the socket, between steps. (The Engine Core Loop) |
| V1 | The current engine architecture, in vllm/v1/. (The Map) |
VllmConfig | The single object carrying every setting to every class. (The Map) |
| Watermark | Fraction of blocks kept free when admitting requests. (Priorities, Preemption and Queues) |
| Worker | The process that owns one GPU. (Executor, Worker, Model Runner) |