The idea in one minute#
Until now the KV cache has lived and died inside one engine. A KV connector is a plug-in point that lets cached blocks leave: to CPU memory or disk when the GPU is full, to another machine that will continue the request, or to a shared store that every replica can read. Three deployment patterns are built on it. Offloading makes the prefix cache larger than the GPU. Prefill/decode disaggregation puts the two halves of a request on different machines so that reading prompts never disturbs streaming. Shared caches let a prefix computed on one replica be used on another. All three rest on the same handful of hooks in the scheduler and the worker, and on one question: is moving these bytes faster than recomputing them?
A picture#
flowchart LR
C[":i-users: Client"] --> PX[":i-route: <b>Proxy / router</b>"]
PX -->|"1. prompt, max_tokens=1"| PRE
PX -->|"3. same prompt + transfer params"| DEC
subgraph PRE["Prefill instance"]
direction TB
P1[":nvidia: <b>vLLM</b><br/><small>kv_role: producer<br/>large token budget</small>"]
end
subgraph DEC["Decode instance"]
direction TB
D1[":nvidia: <b>vLLM</b><br/><small>kv_role: consumer<br/>decode-only CUDA graphs</small>"]
end
P1 ==>|"2. KV blocks<br/>RDMA / NIXL"| D1
D1 -->|"4. streamed tokens"| C
P1 -.-> ST[(":i-database: <b>Offload tiers</b><br/><small>CPU RAM, local disk,<br/>shared store</small>")]
D1 -.-> ST
class C neutral
class PX queue
class P1,D1 compute
class ST memoryHow it really works#
The connector interface#
A connector (KVConnectorBase_V1 in vllm/distributed/kv_transfer/kv_connector/v1/base.py)
has two halves, one in each process that touches the cache.
Scheduler side, in the engine core:
| Hook | Called | Purpose |
|---|---|---|
get_num_new_matched_tokens(request, num_local) | When a waiting request is considered | “Beyond the local cache hit, how many more tokens can you supply?” Also says whether loading is asynchronous. |
update_state_after_alloc(request, blocks, n) | After blocks are allocated | Tells the connector which blocks will receive the loaded data |
build_connector_meta(scheduler_output) | End of schedule() | Packs this step’s load and save instructions for the workers |
request_finished(request, block_ids) | When a request ends | The connector may ask the scheduler to delay freeing the blocks until a transfer completes |
Worker side, in the model runner:
| Hook | Called | Purpose |
|---|---|---|
register_kv_caches(kv_caches) | Startup | Hands the connector the actual cache tensors |
start_load_kv(...) | Before the forward pass | Begin copying external KV into the allocated blocks |
save_kv_layer(...), wait_for_save() | During and after the forward pass | Copy newly computed KV out |
get_finished(...) | Each step | Report which transfers completed |
You have already met every place the scheduler calls these. In
One Scheduling Step the cache lookup for a new
request asks the local prefix cache first and then the connector; in
Allocating Slots the ext_comp segment is the tokens a
connector will supply.
A connector is selected with --kv-transfer-config, a JSON object naming the connector, its
role (kv_producer, kv_consumer or kv_both) and its own settings.
Asynchronous loads and the waiting state#
Fetching a gigabyte of KV from another machine takes time. The engine must not stall for it. When a connector answers “I can supply these tokens, asynchronously”, the scheduler:
if load_kv_async:
# If loading async, allocate memory and put request
# into the WAITING_FOR_REMOTE_KV state.
request.status = RequestStatus.WAITING_FOR_REMOTE_KVS
request.num_computed_tokens = num_computed_tokens
self._inflight_prefills.add(request)
skip_request(request_queue)
continueBlocks are allocated at once, so the destination exists. The request is then set aside, in the
kv_holding_waiting queue, while other requests run. Each step the worker reports finished
transfers; when this one completes, the request becomes schedulable with its
num_computed_tokens already covering the loaded prefix, and it proceeds as if it had a very
large cache hit.
Three protections surround this, all visible in the source:
- Admission is stricter. An in-flight load cannot be preempted and makes no progress of its own, so it is admitted only if it fits in free blocks minus what other in-flight prefills will need: “to avoid deadlock and predictable preemptions.”
- Failures fall back to computing. If a transfer fails, the affected blocks are marked
invalid and
num_computed_tokensis rolled back to the last good position. The request recomputes the rest locally. A request-level failure that cannot be recovered finishes withFinishReason.ERROR, which the API layer always reports as HTTP 500 so it can be retried. - Blocks are not cached until loaded.
delay_cache_blocks=Truekeeps partially loaded blocks out of the local prefix cache.
The loop’s one-millisecond sleep from The Engine Core Loop exists for this case: requests are present, none can run, and the transfer threads need the interpreter.
Pattern 1: offloading#
The simplest use. When a block is evicted from the GPU, its contents are lost. With offloading, completed blocks are also copied to a larger, slower tier, and a later request whose prefix is no longer on the GPU can have it copied back.
vllm serve <model> --kv-offloading-size 64 # GiB of host memory, native backend--kv-offloading-size enables it and sets the buffer size (summed across tensor-parallel
ranks); --kv-offloading-backend chooses native (vLLM’s own) or lmcache. For finer control
the same thing is configured as the OffloadingConnector, with two layouts:
- CPU tier only (default): completed GPU blocks are copied into pinned host memory.
- Tiered: the CPU tier plus secondary tiers such as a filesystem. Only the CPU tier talks to the GPU; everything else is staged through it.
Offloading turns the GPU’s prefix cache into the first level of a hierarchy. It pays off when the set of prefixes worth keeping — many users’ long conversations, a large document library — is bigger than the GPU’s pool.
A request can limit how much it loads with max_load_tokens in kv_transfer_params; zero
disables external loading for that request.
Pattern 2: prefill/decode disaggregation#
Run two pools of vLLM instances. Prefill instances only read prompts; decode instances only generate. The flow for one request:
- A proxy sends the prompt to a prefill instance with
max_tokens: 1and transfer parameters naming the decode instance. - The prefill instance computes the prompt’s KV. Its connector ships the blocks to the decode instance.
- The proxy sends the same request to the decode instance, marked as a remote-prefill request. Its scheduler asks the connector, is told the whole prompt is available, allocates blocks, waits for the transfer, and starts generating.
- Tokens stream from the decode instance to the client.
The documentation gives two reasons to do this, and one warning:
- Tuning time-to-first-token (TTFT) and inter-token-latency (ITL) separately. …
- Controlling tail ITL. Without disaggregated prefilling, vLLM may insert some prefill jobs during the decoding of one request. This results in higher tail latency. …
Disaggregated prefill DOES NOT improve throughput.
That is the trade stated honestly. From Chunked Prefill and the Token Budget: in a shared engine, every prefill chunk lengthens the step for every decoder. Chunked prefill bounds the damage; disaggregation removes it, because decode instances never run a prefill. The documentation notes that chunked prefill “can achieve the same goal, but in practice it’s hard to figure out the correct chunk size value.”
Each pool can then be configured for its job:
| Prefill pool | Decode pool | |
|---|---|---|
| Token budget | Large | Irrelevant; no long prefills |
max_num_seqs | Small | Large |
| CUDA graphs | Piecewise | FULL_DECODE_ONLY, saving the memory of the piecewise graphs |
| Admission | --max-num-queued-tokens sized as target TTFT × prefill throughput | By KV capacity |
| Parallelism | Whatever makes prefill fastest | Whatever gives the lowest step time |
The cost is a second fleet, a proxy that understands the two-step protocol, and a fast network. It is also flagged experimental in the documentation.
Pattern 3: a cache shared between replicas#
With several replicas, a prefix cached on one is useless to a request routed to another. There are two ways to fix that.
Move the request to the cache. Each engine can publish KV events — BlockStored,
BlockRemoved, AllBlocksCleared, with block hashes — to a subscriber. A router that listens
knows which replica holds which prefixes and sends each request where its prefix already is.
Because the default hash is SHA-256 with a fixed seed, the router can compute the same hashes
itself from token IDs (Prefix Caching). This is what the
external load-balancing mode and the tokens-in endpoint exist to support.
Move the cache to the request. A connector backed by a shared store (LMCache, Mooncake store, FlexKV and others) lets any replica fetch blocks another replica computed.
The first costs nothing per request but needs a smart router. The second works with any router but pays a transfer on every cross-replica hit.
The connectors that ship#
| Connector | Transport or store | Typical use |
|---|---|---|
OffloadingConnector | Host memory, optional further tiers | Extending the prefix cache |
NixlConnector | NVIDIA’s NIXL transfer library, over UCX and other backends; fully asynchronous | Prefill/decode disaggregation |
LMCacheConnectorV1, LMCacheMPConnector | LMCache | Shared and offloaded caches; disaggregation |
MooncakeConnector | Mooncake transfer engine and store | Disaggregation; shared store |
MoRIIOConnector | AMD’s MoRI-IO | Disaggregation on ROCm |
FlexKVConnectorV1 | FlexKV distributed store | Multi-level cache at very large scale |
MultiConnector | Several of the above, in priority order | For example NIXL for hand-off plus a store for reuse |
ExampleConnector | Files on shared storage | Learning and testing only |
Is transferring worth it?#
A connector helps only if moving the KV is faster than recomputing it.
KV bytes = tokens × bytes per token 8,000 × 128 KiB ≈ 1 GiB for an 8B model
transfer = bytes / bandwidth
recompute = tokens / prefill speedRecomputing gets relatively cheaper as models get smaller and GPUs faster; transferring gets cheaper as networks get faster and caches get quantised. For a small model on a 10-gigabit network, recomputing wins. For a large model with RDMA it is not close, in the other direction. The program below works the numbers.
Two requirements are easy to overlook:
- Both ends must agree exactly. Same model, same KV cache data type, same block size, same attention layout. Connectors exchange handshake metadata at startup to check.
- The prefill instance must release its blocks only after sending. That is the
request_finishedhook’s “delay free” return value, and the reason the scheduler has a deferred-free path at all.
Code#
When does moving KV beat recomputing it? Vary the model, the network and the cache precision.
package main
import "fmt"
func main() {
type model struct {
name string
kvKiBPerTok float64 // 16-bit cache
prefillTokPS float64 // prompt tokens per second on one GPU
}
type link struct {
name string
gbitPS float64
}
models := []model{
{"1.5B model", 28, 60000},
{"8B model", 128, 20000},
{"70B model (tp 4)", 320, 6000},
}
links := []link{
{"10 GbE", 10},
{"100 GbE / RDMA", 100},
{"host RAM (offload)", 800}, // a memory copy over PCIe, order of magnitude
}
const promptTokens = 8000
for _, kvBits := range []float64{16, 8} {
fmt.Printf("%d-token prompt, %d-bit KV cache\n", promptTokens, int(kvBits))
fmt.Printf(" %-18s %10s %10s", "model", "KV size", "recompute")
for _, l := range links {
fmt.Printf(" %20s", l.name)
}
fmt.Println()
for _, m := range models {
bytes := promptTokens * m.kvKiBPerTok * 1024 * kvBits / 16
recompute := promptTokens / m.prefillTokPS
fmt.Printf(" %-18s %7.2f GiB %8.2f s", m.name, bytes/(1<<30), recompute)
for _, l := range links {
transfer := bytes * 8 / (l.gbitPS * 1e9)
verdict := "transfer"
if transfer > recompute {
verdict = "recompute"
}
fmt.Printf(" %9.2f s %-10s", transfer, verdict)
}
fmt.Println()
}
fmt.Println()
}
}Read across a row: the link decides. On 10-gigabit Ethernet with a 16-bit cache, recomputing wins for every model here; on a 100-gigabit link, or for a copy to host memory, transferring wins for every model by a wide margin. Compare the two tables: an 8-bit cache halves every transfer time, which is enough to flip some of the 10-gigabit rows. The practical reading is that offloading to host memory is nearly always worthwhile, and moving KV between machines is worthwhile only on a fast network — which is why none of this machinery is on by default.
Remember this#
- A KV connector lets cached blocks leave the engine: to other tiers, other machines, or a shared store.
- Scheduler-side hooks answer “how many more tokens can you supply?”; worker-side hooks move the bytes.
- Asynchronous loads put a request in
WAITING_FOR_REMOTE_KVSwith blocks already allocated; failures fall back to recomputing. - Offloading extends the prefix cache beyond GPU memory:
--kv-offloading-size. - Prefill/decode disaggregation isolates decoders from prefills. It improves tail latency, not throughput.
- Replicas can share cache by routing requests to it (KV events) or by fetching it (a shared store).
- It is worth it only when transfer time is below recompute time: large models, fast links, quantised caches.
Try it#
- Change
promptTokensto 100,000. Which cells flip, and why does the answer not depend on prompt length in this simple model? What real-world effects would make it depend on length? - Add a “local NVMe” link at 20 Gbit/s. For which models is a disk tier worth having?
- Start a server with
--kv-offloading-size 8, fill the GPU cache with many distinct long prompts, then resend the first one. Compare its time to first token with and without offloading.
Check yourself#
- Which two processes does a KV connector have code in, and what does each half do?
- Why does prefill/decode disaggregation not increase throughput?
- What must be identical between a prefill instance and a decode instance for a transfer to be valid?
Sources#
Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.
- Disaggregated prefilling
- KV offloading usage guide
- NixlConnector usage guide
vllm/distributed/kv_transfer/kv_connector/v1/base.py— the connector interfacevllm/v1/core/sched/scheduler.py—WAITING_FOR_REMOTE_KVS, failed loads, deferred freesvllm/config/cache.py—kv_offloading_size,kv_offloading_backend