The idea in one minute#
A fine-tuned model is usually the base model plus a small correction. LoRA stores that correction as two thin matrices per layer — a few megabytes instead of many gigabytes. vLLM can hold dozens of such adapters next to one copy of the base weights and apply a different one to each request in the same batch. To the client each adapter looks like a separate model name; on the GPU it is one model whose linear layers add a per-row correction. The scheduler, the prefix cache and the CUDA graph dispatcher each needed one small change to make that safe.
A picture#
flowchart LR
subgraph BATCH["One batch"]
direction TB
R1["request 1<br/><small>model: sql-lora</small>"]
R2["request 2<br/><small>model: base</small>"]
R3["request 3<br/><small>model: support-lora</small>"]
end
BATCH --> L[":pytorch: <b>Linear layer with LoRA</b>"]
W[(":i-hard-drive: <b>Base weight W</b><br/><small>one copy, shared</small>")] --> L
A[(":i-layers: <b>Adapter slots</b><br/><small>stacked A and B matrices<br/>max_loras slots</small>")] --> L
L --> O["y = W·x + B<sub>i</sub>·A<sub>i</sub>·x<br/><small>i = the row's adapter, or none</small>"]
CPU[(":i-memory-stick: <b>CPU adapter cache</b><br/><small>LRU, max_cpu_loras</small>")] -.->|"activate into a slot"| A
class R1,R2,R3 neutral
class L,O compute
class W,A,CPU memoryHow it really works#
What an adapter is#
For a layer with weight W of size d × d, LoRA trains two matrices A (r × d) and B
(d × r) with a small rank r, typically 8 to 64. The adapted layer computes
y = W·x + (α / r) · B·A·xW is untouched. With d = 4096 and r = 16, A and B together hold 131,072 numbers
against W’s 16.8 million: under 1%.
There are two ways to serve this. Merging computes W' = W + B·A once and serves an
ordinary model: zero overhead, but one full copy of the weights per adapter. vLLM instead keeps
W shared and adds the correction at run time, so one base model serves many adapters.
Using it#
vllm serve meta-llama/Llama-3.2-3B-Instruct \
--enable-lora \
--lora-modules sql-lora=jeeejeee/llama32-3b-text2sql-spider \
--max-loras 4 --max-lora-rank 32Each adapter appears in /v1/models as its own model. A client selects one by putting its
name in the request’s model field; the base model’s name selects no adapter.
| Flag | Default | Meaning |
|---|---|---|
--enable-lora | off | Build the model with LoRA-capable layers |
--max-loras | 1 | Distinct adapters allowed in one batch; also the number of GPU slots |
--max-lora-rank | 16 | Largest rank any adapter may have. Slots are sized for this. |
--max-cpu-loras | = max_loras | Adapters kept loaded in CPU memory, least-recently-used eviction |
--lora-modules | — | Adapters to register at startup, name=path |
Setting --max-lora-rank above what you need wastes memory and compute: every slot is
allocated at the maximum rank and smaller adapters are padded.
Three tiers of storage#
| Tier | Holds | Limit | Cost to reach the next tier |
|---|---|---|---|
| Disk or hub | Every adapter you could load | Unlimited | Download and parse: slow |
| CPU cache | Parsed adapter tensors | max_cpu_loras | Copy into a GPU slot: fast |
| GPU slots | Adapters usable in this step | max_loras | — |
An adapter in the CPU cache can be activated into a GPU slot in milliseconds, so a server
with --max-loras 4 --max-cpu-loras 64 can serve 64 adapters, four at a time per batch.
The scheduler’s part#
The scheduler enforces the batch limit. While building a step it tracks which adapters are already in it:
# Check that adding the request still respects the max_loras constraint.
if (
self.lora_config
and request.lora_request
and (
len(scheduled_loras) == self.lora_config.max_loras
and request.lora_request.lora_int_id not in scheduled_loras
)
):
# Scheduling would exceed max_loras, skip.
skip_request(request_queue)
continueA waiting request whose adapter would be the (max_loras + 1)-th distinct adapter is
skipped, not blocked on: requests behind it that use an adapter already in the batch, or
none, can still be admitted. It returns to the front of the queue for the next step.
This is the main operational risk of multi-adapter serving. With max_loras = 1 and traffic
spread evenly over ten adapters, only one adapter’s requests run at a time; the other nine
wait until the running ones finish. Throughput collapses although the GPU is not full. Size
max_loras to the number of adapters that are simultaneously active, not the number that
exist.
The prefix cache’s part#
The same prompt produces different KV under different adapters, so the adapter is part of every block hash:
def _gen_lora_extra_hash_keys(request: Request) -> list[tuple[str, str, str]]:
"""...
The adapter path is included so that re-pointing a LoRA name at a different
adapter does not reuse KV computed with the previous one.
"""
lora_request = request.lora_request
if not lora_request:
return []
return [("lora", lora_request.lora_name, lora_request.lora_path)]Requests using adapter A never share cached blocks with requests using adapter B or the base model, even for an identical system prompt. Ten adapters with a common 4,000-token system prompt hold ten copies of its KV. Budget cache memory accordingly.
The model runner’s part#
Each LoRA-capable layer keeps stacked A and B tensors with one slot per active adapter,
and receives a per-token index saying which slot (or none) applies to each row of the flat
batch. The correction is computed for all rows in one batched operation.
Two details from the configuration show where the cost goes:
max_num_batched_tokensandmax_num_seqsare part of the compilation cache key because, in the words ofSchedulerConfig.compute_hash, “LoRA creates static buffers based onmax_num_batched_tokens. The tensor sizes and strides get captured in thetorch.compilegraph explicitly.”- A batch with adapters active executes different operations from one without, so the CUDA
graph dispatcher keys on
has_lora(CUDA Graphs and Compilation). Enabling LoRA roughly doubles the number of graphs captured. An option (specialize_active_lora) captures separate graphs per number of active adapters, trading more startup time and memory for speed when the count varies.
Loading adapters while running#
With VLLM_ALLOW_RUNTIME_LORA_UPDATING=True, two endpoints manage adapters without a restart:
curl -X POST http://localhost:8000/v1/load_lora_adapter \
-H "Content-Type: application/json" \
-d '{"lora_name": "support-lora", "lora_path": "/adapters/support-v3"}'
curl -X POST http://localhost:8000/v1/unload_lora_adapter \
-H "Content-Type: application/json" \
-d '{"lora_name": "support-lora"}'Internally these are utility calls executed between engine steps (The Engine Core Loop), which is why no lock is needed around the adapter tables.
warning
Runtime loading lets anyone who can reach the endpoint make the server read an arbitrary path or download an arbitrary repository. The flag is off by default for that reason. In a replicated deployment it also creates drift: an adapter loaded on one replica does not exist on the others. The documentation recommends runtime loading for single-node use and a resolver plugin for production.
A LoRA resolver plugin inverts the flow: when a request names an unknown model, vLLM asks the registered resolvers whether they can provide an adapter of that name, and loads it on demand. Every replica resolves the same name the same way, so there is no drift (Plugins).
What it costs#
- Per token: the extra low-rank multiply in every adapted layer. A few percent to low tens of percent of decode time, depending on rank, the number of adapted modules and batch mix.
- Memory:
max_loras × layers × modules × 2 × d × max_lora_ranknumbers. Small next to the base model, not negligible at high rank and many slots. - Cache efficiency: no prefix sharing across adapters.
- Scheduling: head-of-line effects when distinct adapters exceed
max_loras.
When one adapter carries nearly all the traffic, merging it into the base weights and serving the result as a plain model removes all four costs.
Code#
Adapter memory, and the scheduling effect of max_loras: the same traffic over eight adapters
with different slot counts.
package main
import "fmt"
func main() {
// --- memory ---
const (
d = 4096 // hidden size
layers = 32
modules = 4 // adapted projections per layer, e.g. q, k, v, o
)
base := float64(layers*modules) * float64(d*d) * 2 / (1 << 30) // GiB at 16 bits
fmt.Printf("base weights in the adapted projections: %.1f GiB\n\n", base)
fmt.Printf("%6s %10s %22s\n", "rank", "one adapter", "GPU slots for max_loras=8")
for _, r := range []int{8, 16, 64, 256} {
one := float64(layers*modules*2*d*r) * 2 / (1 << 20) // MiB
fmt.Printf("%6d %8.0f MiB %18.2f GiB\n", r, one, one*8/1024)
}
// --- scheduling ---
fmt.Println("\n64 requests, 8 per adapter, arriving interleaved; each needs 20 decode steps.")
fmt.Printf("%10s %8s %22s\n", "max_loras", "steps", "avg requests per step")
for _, maxLoras := range []int{1, 2, 4, 8} {
type req struct{ adapter, left int }
var waiting []*req
for i := 0; i < 64; i++ {
waiting = append(waiting, &req{adapter: i % 8, left: 20})
}
var running []*req
steps, served := 0, 0
for len(waiting)+len(running) > 0 {
// Adapters already in the batch.
active := map[int]bool{}
for _, r := range running {
active[r.adapter] = true
}
// Admit from the queue; skip requests whose adapter would exceed max_loras.
var still []*req
for _, r := range waiting {
if active[r.adapter] || len(active) < maxLoras {
active[r.adapter] = true
running = append(running, r)
} else {
still = append(still, r) // skipped, keeps its place
}
}
waiting = still
steps++
served += len(running)
keep := running[:0]
for _, r := range running {
if r.left--; r.left > 0 {
keep = append(keep, r)
}
}
running = keep
}
fmt.Printf("%10d %8d %22.1f\n", maxLoras, steps, float64(served)/float64(steps))
}
}With one slot the server takes eight times as many steps and runs batches of eight; with eight slots all 64 requests share every step. The GPU, the model and the traffic are identical in both rows. Only the flag changed.
Remember this#
- vLLM adds the low-rank correction at run time, so one copy of the base weights serves many adapters.
- An adapter is selected by using its name as the request’s
model. --max-lorasis the number of distinct adapters per batch, not the number you can serve.- A request whose adapter would exceed
max_lorasis skipped until a slot frees. - The adapter is part of every prefix-cache hash; adapters never share cached blocks.
- LoRA adds a
has_loradimension to CUDA graphs and static buffers tied to the batch limits. - Runtime loading is off by default; use a resolver plugin for replicated deployments.
Try it#
- In the scheduling experiment, make the traffic skewed: 50 requests for adapter 0 and 2 each
for the other seven. How much does
max_loras = 2cost now compared with 8? - Compute the slot memory for your own model and a rank of 64 with 16 slots.
- On a real server with two adapters and
--max-loras 1, send alternating requests and watchvllm:num_requests_waiting. Repeat with--max-loras 2.
Check yourself#
- Why does vLLM not merge adapters into the base weights?
- What happens to a waiting request whose adapter is not in the batch when the batch already holds
max_lorasdistinct adapters? - Two tenants use different adapters and the same system prompt. Do they share prefix-cache blocks? Why?
Sources#
Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.
- LoRA adapters
vllm/config/lora.pyvllm/v1/core/sched/scheduler.py— themax_lorascheckvllm/v1/core/kv_cache_utils.py—_gen_lora_extra_hash_keys- LoRA resolver plugins
- LoRA: Low-Rank Adaptation of Large Language Models