Pidoku

Disaggregation and KV Connectors

Expert 50 min Difficulty 4/5 Lesson 02 of 05

Prerequisites Prefix Caching, Allocating Slots, Parallelism

The idea in one minute#

Until now the KV cache has lived and died inside one engine. A KV connector is a plug-in point that lets cached blocks leave: to CPU memory or disk when the GPU is full, to another machine that will continue the request, or to a shared store that every replica can read. Three deployment patterns are built on it. Offloading makes the prefix cache larger than the GPU. Prefill/decode disaggregation puts the two halves of a request on different machines so that reading prompts never disturbs streaming. Shared caches let a prefix computed on one replica be used on another. All three rest on the same handful of hooks in the scheduler and the worker, and on one question: is moving these bytes faster than recomputing them?

A picture#

flowchart LR
  C[":i-users: Client"] --> PX[":i-route: <b>Proxy / router</b>"]
  PX -->|"1. prompt, max_tokens=1"| PRE
  PX -->|"3. same prompt + transfer params"| DEC
  subgraph PRE["Prefill instance"]
    direction TB
    P1[":nvidia: <b>vLLM</b><br/><small>kv_role: producer<br/>large token budget</small>"]
  end
  subgraph DEC["Decode instance"]
    direction TB
    D1[":nvidia: <b>vLLM</b><br/><small>kv_role: consumer<br/>decode-only CUDA graphs</small>"]
  end
  P1 ==>|"2. KV blocks<br/>RDMA / NIXL"| D1
  D1 -->|"4. streamed tokens"| C
  P1 -.-> ST[(":i-database: <b>Offload tiers</b><br/><small>CPU RAM, local disk,<br/>shared store</small>")]
  D1 -.-> ST
  class C neutral
  class PX queue
  class P1,D1 compute
  class ST memory

How it really works#

The connector interface#

A connector (KVConnectorBase_V1 in vllm/distributed/kv_transfer/kv_connector/v1/base.py) has two halves, one in each process that touches the cache.

Scheduler side, in the engine core:

HookCalledPurpose
get_num_new_matched_tokens(request, num_local)When a waiting request is considered“Beyond the local cache hit, how many more tokens can you supply?” Also says whether loading is asynchronous.
update_state_after_alloc(request, blocks, n)After blocks are allocatedTells the connector which blocks will receive the loaded data
build_connector_meta(scheduler_output)End of schedule()Packs this step’s load and save instructions for the workers
request_finished(request, block_ids)When a request endsThe connector may ask the scheduler to delay freeing the blocks until a transfer completes

Worker side, in the model runner:

HookCalledPurpose
register_kv_caches(kv_caches)StartupHands the connector the actual cache tensors
start_load_kv(...)Before the forward passBegin copying external KV into the allocated blocks
save_kv_layer(...), wait_for_save()During and after the forward passCopy newly computed KV out
get_finished(...)Each stepReport which transfers completed

You have already met every place the scheduler calls these. In One Scheduling Step the cache lookup for a new request asks the local prefix cache first and then the connector; in Allocating Slots the ext_comp segment is the tokens a connector will supply.

A connector is selected with --kv-transfer-config, a JSON object naming the connector, its role (kv_producer, kv_consumer or kv_both) and its own settings.

Asynchronous loads and the waiting state#

Fetching a gigabyte of KV from another machine takes time. The engine must not stall for it. When a connector answers “I can supply these tokens, asynchronously”, the scheduler:

Python
if load_kv_async:
    # If loading async, allocate memory and put request
    # into the WAITING_FOR_REMOTE_KV state.
    request.status = RequestStatus.WAITING_FOR_REMOTE_KVS
    request.num_computed_tokens = num_computed_tokens
    self._inflight_prefills.add(request)
    skip_request(request_queue)
    continue

Blocks are allocated at once, so the destination exists. The request is then set aside, in the kv_holding_waiting queue, while other requests run. Each step the worker reports finished transfers; when this one completes, the request becomes schedulable with its num_computed_tokens already covering the loaded prefix, and it proceeds as if it had a very large cache hit.

Three protections surround this, all visible in the source:

  • Admission is stricter. An in-flight load cannot be preempted and makes no progress of its own, so it is admitted only if it fits in free blocks minus what other in-flight prefills will need: “to avoid deadlock and predictable preemptions.”
  • Failures fall back to computing. If a transfer fails, the affected blocks are marked invalid and num_computed_tokens is rolled back to the last good position. The request recomputes the rest locally. A request-level failure that cannot be recovered finishes with FinishReason.ERROR, which the API layer always reports as HTTP 500 so it can be retried.
  • Blocks are not cached until loaded. delay_cache_blocks=True keeps partially loaded blocks out of the local prefix cache.

The loop’s one-millisecond sleep from The Engine Core Loop exists for this case: requests are present, none can run, and the transfer threads need the interpreter.

Pattern 1: offloading#

The simplest use. When a block is evicted from the GPU, its contents are lost. With offloading, completed blocks are also copied to a larger, slower tier, and a later request whose prefix is no longer on the GPU can have it copied back.

Shell
vllm serve <model> --kv-offloading-size 64          # GiB of host memory, native backend

--kv-offloading-size enables it and sets the buffer size (summed across tensor-parallel ranks); --kv-offloading-backend chooses native (vLLM’s own) or lmcache. For finer control the same thing is configured as the OffloadingConnector, with two layouts:

  • CPU tier only (default): completed GPU blocks are copied into pinned host memory.
  • Tiered: the CPU tier plus secondary tiers such as a filesystem. Only the CPU tier talks to the GPU; everything else is staged through it.

Offloading turns the GPU’s prefix cache into the first level of a hierarchy. It pays off when the set of prefixes worth keeping — many users’ long conversations, a large document library — is bigger than the GPU’s pool.

A request can limit how much it loads with max_load_tokens in kv_transfer_params; zero disables external loading for that request.

Pattern 2: prefill/decode disaggregation#

Run two pools of vLLM instances. Prefill instances only read prompts; decode instances only generate. The flow for one request:

  1. A proxy sends the prompt to a prefill instance with max_tokens: 1 and transfer parameters naming the decode instance.
  2. The prefill instance computes the prompt’s KV. Its connector ships the blocks to the decode instance.
  3. The proxy sends the same request to the decode instance, marked as a remote-prefill request. Its scheduler asks the connector, is told the whole prompt is available, allocates blocks, waits for the transfer, and starts generating.
  4. Tokens stream from the decode instance to the client.

The documentation gives two reasons to do this, and one warning:

  • Tuning time-to-first-token (TTFT) and inter-token-latency (ITL) separately. …
  • Controlling tail ITL. Without disaggregated prefilling, vLLM may insert some prefill jobs during the decoding of one request. This results in higher tail latency. …

Disaggregated prefill DOES NOT improve throughput.

That is the trade stated honestly. From Chunked Prefill and the Token Budget: in a shared engine, every prefill chunk lengthens the step for every decoder. Chunked prefill bounds the damage; disaggregation removes it, because decode instances never run a prefill. The documentation notes that chunked prefill “can achieve the same goal, but in practice it’s hard to figure out the correct chunk size value.”

Each pool can then be configured for its job:

Prefill poolDecode pool
Token budgetLargeIrrelevant; no long prefills
max_num_seqsSmallLarge
CUDA graphsPiecewiseFULL_DECODE_ONLY, saving the memory of the piecewise graphs
Admission--max-num-queued-tokens sized as target TTFT × prefill throughputBy KV capacity
ParallelismWhatever makes prefill fastestWhatever gives the lowest step time

The cost is a second fleet, a proxy that understands the two-step protocol, and a fast network. It is also flagged experimental in the documentation.

Pattern 3: a cache shared between replicas#

With several replicas, a prefix cached on one is useless to a request routed to another. There are two ways to fix that.

Move the request to the cache. Each engine can publish KV events — BlockStored, BlockRemoved, AllBlocksCleared, with block hashes — to a subscriber. A router that listens knows which replica holds which prefixes and sends each request where its prefix already is. Because the default hash is SHA-256 with a fixed seed, the router can compute the same hashes itself from token IDs (Prefix Caching). This is what the external load-balancing mode and the tokens-in endpoint exist to support.

Move the cache to the request. A connector backed by a shared store (LMCache, Mooncake store, FlexKV and others) lets any replica fetch blocks another replica computed.

The first costs nothing per request but needs a smart router. The second works with any router but pays a transfer on every cross-replica hit.

The connectors that ship#

ConnectorTransport or storeTypical use
OffloadingConnectorHost memory, optional further tiersExtending the prefix cache
NixlConnectorNVIDIA’s NIXL transfer library, over UCX and other backends; fully asynchronousPrefill/decode disaggregation
LMCacheConnectorV1, LMCacheMPConnectorLMCacheShared and offloaded caches; disaggregation
MooncakeConnectorMooncake transfer engine and storeDisaggregation; shared store
MoRIIOConnectorAMD’s MoRI-IODisaggregation on ROCm
FlexKVConnectorV1FlexKV distributed storeMulti-level cache at very large scale
MultiConnectorSeveral of the above, in priority orderFor example NIXL for hand-off plus a store for reuse
ExampleConnectorFiles on shared storageLearning and testing only

Is transferring worth it?#

A connector helps only if moving the KV is faster than recomputing it.

KV bytes  = tokens × bytes per token          8,000 × 128 KiB ≈ 1 GiB for an 8B model
transfer  = bytes / bandwidth
recompute = tokens / prefill speed

Recomputing gets relatively cheaper as models get smaller and GPUs faster; transferring gets cheaper as networks get faster and caches get quantised. For a small model on a 10-gigabit network, recomputing wins. For a large model with RDMA it is not close, in the other direction. The program below works the numbers.

Two requirements are easy to overlook:

  • Both ends must agree exactly. Same model, same KV cache data type, same block size, same attention layout. Connectors exchange handshake metadata at startup to check.
  • The prefill instance must release its blocks only after sending. That is the request_finished hook’s “delay free” return value, and the reason the scheduler has a deferred-free path at all.

Code#

When does moving KV beat recomputing it? Vary the model, the network and the cache precision.

Go
package main

import "fmt"

func main() {
	type model struct {
		name         string
		kvKiBPerTok  float64 // 16-bit cache
		prefillTokPS float64 // prompt tokens per second on one GPU
	}
	type link struct {
		name   string
		gbitPS float64
	}
	models := []model{
		{"1.5B model", 28, 60000},
		{"8B model", 128, 20000},
		{"70B model (tp 4)", 320, 6000},
	}
	links := []link{
		{"10 GbE", 10},
		{"100 GbE / RDMA", 100},
		{"host RAM (offload)", 800}, // a memory copy over PCIe, order of magnitude
	}
	const promptTokens = 8000

	for _, kvBits := range []float64{16, 8} {
		fmt.Printf("%d-token prompt, %d-bit KV cache\n", promptTokens, int(kvBits))
		fmt.Printf("  %-18s %10s %10s", "model", "KV size", "recompute")
		for _, l := range links {
			fmt.Printf(" %20s", l.name)
		}
		fmt.Println()
		for _, m := range models {
			bytes := promptTokens * m.kvKiBPerTok * 1024 * kvBits / 16
			recompute := promptTokens / m.prefillTokPS
			fmt.Printf("  %-18s %7.2f GiB %8.2f s", m.name, bytes/(1<<30), recompute)
			for _, l := range links {
				transfer := bytes * 8 / (l.gbitPS * 1e9)
				verdict := "transfer"
				if transfer > recompute {
					verdict = "recompute"
				}
				fmt.Printf(" %9.2f s %-10s", transfer, verdict)
			}
			fmt.Println()
		}
		fmt.Println()
	}
}

Read across a row: the link decides. On 10-gigabit Ethernet with a 16-bit cache, recomputing wins for every model here; on a 100-gigabit link, or for a copy to host memory, transferring wins for every model by a wide margin. Compare the two tables: an 8-bit cache halves every transfer time, which is enough to flip some of the 10-gigabit rows. The practical reading is that offloading to host memory is nearly always worthwhile, and moving KV between machines is worthwhile only on a fast network — which is why none of this machinery is on by default.

Remember this#

  • A KV connector lets cached blocks leave the engine: to other tiers, other machines, or a shared store.
  • Scheduler-side hooks answer “how many more tokens can you supply?”; worker-side hooks move the bytes.
  • Asynchronous loads put a request in WAITING_FOR_REMOTE_KVS with blocks already allocated; failures fall back to recomputing.
  • Offloading extends the prefix cache beyond GPU memory: --kv-offloading-size.
  • Prefill/decode disaggregation isolates decoders from prefills. It improves tail latency, not throughput.
  • Replicas can share cache by routing requests to it (KV events) or by fetching it (a shared store).
  • It is worth it only when transfer time is below recompute time: large models, fast links, quantised caches.

Try it#

  1. Change promptTokens to 100,000. Which cells flip, and why does the answer not depend on prompt length in this simple model? What real-world effects would make it depend on length?
  2. Add a “local NVMe” link at 20 Gbit/s. For which models is a disk tier worth having?
  3. Start a server with --kv-offloading-size 8, fill the GPU cache with many distinct long prompts, then resend the first one. Compare its time to first token with and without offloading.

Check yourself#

  1. Which two processes does a KV connector have code in, and what does each half do?
  2. Why does prefill/decode disaggregation not increase throughput?
  3. What must be identical between a prefill instance and a decode instance for a transfer to be valid?

Sources#

Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom