Pidoku
Glossary

Glossary

Short definitions of the terms this course uses. The lesson in brackets is where each is explained.

TermMeaning
Admission controlRefusing new requests with HTTP 503 when queue limits are reached. (The Frontend)
All-reduceA collective operation that sums partial results across GPUs; two per layer under tensor parallelism. (Parallelism)
APCAutomatic prefix caching. (Prefix Caching)
API serverThe process that handles HTTP, templates, tokenising and detokenising. (Processes and Wires)
Argmax-invariantA logits transformation that cannot change which token scores highest. (Sampling and Structured Output)
Async schedulingScheduling step N+1 while step N is still on the GPU. (Async Scheduling)
AsyncLLMThe engine as seen from an asyncio program; what the server uses. (AsyncLLM and the Way Back)
Attention backendA kernel implementation that can read a paged KV cache, chosen at startup. (Attention Backends)
Batch queueBatches scheduled but not yet finished; size 2 with async scheduling. (The Engine Core Loop)
BlockFixed-size unit of KV cache: 16 tokens by default. (Blocks and the Pool)
Block hashA hash of a block’s tokens and its parent’s hash; identifies a whole prefix. (Prefix Caching)
Block poolAll KV blocks, the free queue and the hash-to-block map. (Blocks and the Pool)
Block tableA request’s ordered list of physical block IDs. (Blocks and the Pool)
Busy loopThe engine core’s main loop; never sleeps while requests exist. (The Engine Core Loop)
cache_saltA per-request value hashed into the first block to isolate tenants’ caches. (Prefix Caching)
Cascade attentionComputing attention over a prefix shared by all requests once. (Attention Backends)
Chat templateA Jinja2 program that renders messages and tools into the model’s prompt format. (The Frontend)
Chunked prefillReading a long prompt over several steps, bounded by the token budget. (Chunked Prefill and the Token Budget)
Closed-loop loadA benchmark where a new request is sent only when one finishes; hides overload. (Tuning and Benchmarking)
Collective RPCCalling a method on every worker and gathering results. (Executor, Worker, Model Runner)
Constrained decodingMasking tokens a grammar forbids before sampling. (Sampling and Structured Output)
Continuous batchingRe-deciding the batch every step so requests join and leave freely. (What vLLM Is and Why It Exists)
Copy-on-writeCopying a shared, partly matching block into a private one before writing. (Prefix Caching)
CUDA graphA recording of GPU operations for one batch shape, replayed in one launch. (CUDA Graphs and Compilation)
Data parallelism (DP)Several complete engines behind one endpoint. (Parallelism)
Deferred freeHolding a finished request’s blocks until in-flight steps that may write them complete. (Async Scheduling)
Detokenizer, incrementalConverts token IDs to text step by step, emitting only complete characters. (AsyncLLM and the Way Back)
Disaggregation (P/D)Running prefill and decode on different instances. (Disaggregation and KV Connectors)
DP coordinatorProcess that balances load and aligns steps across data-parallel ranks. (Parallelism)
Encoder cacheHolds a vision or audio encoder’s output until the request has consumed it. (Multimodal Inputs)
Engine coreThe process that runs the scheduler and KV cache manager. (Processes and Wires)
EngineCoreRequest / EngineCoreOutputThe msgpack messages between API server and engine core. (Processes and Wires)
EvictionA cached block losing its hash when handed to a new owner; lazy in vLLM. (Blocks and the Pool)
ExecutorDelivers one call to every worker of an engine. (Executor, Worker, Model Runner)
Expert parallelism (EP)Distributing a mixture-of-experts model’s experts across ranks. (Parallelism)
Flat batchAll scheduled tokens of all requests in one row, with index tensors instead of padding. (Building the Batch)
Free queueDoubly linked list of unowned blocks; its order is the eviction policy. (Blocks and the Pool)
GapTokens a request has minus tokens computed; what the scheduler closes. (The Map)
Grammar bitmaskOne bit per vocabulary token saying whether it is legal next. (Sampling and Structured Output)
Head-of-line blockingA waiting request that does not fit stops admission behind it. (One Scheduling Step)
Inter-token latency (ITL)Time between consecutive streamed outputs. (Metrics)
KV cache groupLayers with the same cache needs, sharing one block table per request. (More Than One Kind of Cache)
KV cache specA layer’s declaration of what its cache stores and how large a page is. (Sizing the Cache)
KV connectorPlug-in that moves cached blocks to other tiers or machines. (Disaggregation and KV Connectors)
KV eventsPublished notices of blocks stored and removed, for cache-aware routers. (Disaggregation and KV Connectors)
logits_indicesRows of the output that become next-token distributions: each request’s last. (Building the Batch)
LookaheadExtra cache slots reserved for a speculative proposer. (Allocating Slots)
LoRAA low-rank correction to a base model, applied per request at run time. (LoRA Adapters)
max_num_batched_tokensThe token budget: most tokens processed in one step. (Chunked Prefill and the Token Budget)
max_num_seqsMost requests running at once, per engine. (One Scheduling Step)
Model runnerBuilds the batch, runs the model and samples, inside a worker. (Executor, Worker, Model Runner)
MRV2Model Runner V2: the newer runner with fixed rows and GPU-side input preparation. (Executor, Worker, Model Runner)
Null blockBlock 0; a placeholder never allocated or freed. (Blocks and the Pool)
num_computed_tokensPositions whose KV is computed or scheduled to be. (One Scheduling Step)
OffloadingCopying cached blocks to CPU memory or disk to extend the prefix cache. (Disaggregation and KV Connectors)
Open-loop loadA benchmark with arrivals at a fixed rate regardless of server state. (Tuning and Benchmarking)
Optimisation level-O0 to -O3: presets trading startup time for speed. (CUDA Graphs and Compilation)
Output placeholderA scheduled position whose token ID is not known yet. (Async Scheduling)
Page sizeBytes one block occupies for one layer or group. (Sizing the Cache)
PagedAttentionAttention over a KV cache stored in scattered fixed-size blocks. (What vLLM Is and Why It Exists)
Persistent batchPer-request state kept between steps and updated by differences. (Executor, Worker, Model Runner)
Piecewise compilationSplitting the model’s graph at attention calls so the rest can be recorded. (CUDA Graphs and Compilation)
Pipeline parallelism (PP)Splitting the layer stack into stages on different GPUs. (Parallelism)
Placeholder tokenA token standing in for one vector of an image or audio clip. (Multimodal Inputs)
PluginCode discovered through Python entry points and loaded in every vLLM process. (Plugins and Extension Points)
PreemptionTaking all of a running request’s blocks; it recomputes when re-admitted. (Priorities, Preemption and Queues)
Prefix cachingReusing KV blocks for a prompt prefix seen before. (Prefix Caching)
prefix_match_unitHashing granularity, which may be finer than the physical block. (Prefix Caching)
query_start_locOffsets marking each request’s rows in the flat batch. (Building the Batch)
Reasoning parserSplits a model’s thinking from its answer into a separate field. (Tool Calling and Reasoning)
Reference countNumber of requests whose block table contains a block. (Blocks and the Pool)
RendererApplies the chat template and tokenises. (The Frontend)
Request statusWAITING, RUNNING, PREEMPTED, several waiting-for states, and the finished states. (One Scheduling Step)
Ring buffer, shared-memoryThe lock-free broadcast channel from engine core to workers. (Processes and Wires)
SchedulerOutputOne step’s plan: which requests, how many tokens, which blocks. (One Scheduling Step)
Sleep modeReleasing GPU memory while keeping the process alive. (The Engine Core Loop)
Slot mappingFor each new token, the physical cache slot its key and value are written to. (Building the Batch)
Speculative decodingProposing several tokens cheaply and verifying them in one pass. (Speculative Decoding)
Stale outputA result from a step scheduled before a request was preempted. (Async Scheduling)
StepOne schedule-execute-update cycle; at most one token per request without speculation. (The Engine Core Loop)
Stop stringText that ends generation; checked in the API server, not the engine. (AsyncLLM and the Way Back)
Structured outputOutput guaranteed to match a JSON schema, regex or grammar. (Sampling and Structured Output)
Tensor parallelism (TP)Splitting each layer’s weights across GPUs. (Parallelism)
Time per output token (TPOT)Per request: (end-to-end latency − TTFT) ÷ (output tokens − 1). (Metrics)
Time to first token (TTFT)Arrival to first streamed token. (Metrics)
Token budgetSee max_num_batched_tokens.
Tool parserExtracts tool calls from a model’s free-text output. (Tool Calling and Reasoning)
TouchA cache hit on a block: reference count up, removed from the free queue. (Blocks and the Pool)
Uniform batchEvery request contributing the same number of tokens; eligible for a full CUDA graph. (CUDA Graphs and Compilation)
Utility callA method invoked on the engine core over the socket, between steps. (The Engine Core Loop)
V1The current engine architecture, in vllm/v1/. (The Map)
VllmConfigThe single object carrying every setting to every class. (The Map)
WatermarkFraction of blocks kept free when admitting requests. (Priorities, Preemption and Queues)
WorkerThe process that owns one GPU. (Executor, Worker, Model Runner)

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom