Pidoku

Adding a Model

Expert 40 min Difficulty 4/5 Lesson 02 of 03

Prerequisites Executor, Worker, Model Runner; Building the Batch; Plugins and Extension Points

The idea in one minute#

Running a new architecture in vLLM is mostly subtraction. You start from the model’s reference PyTorch code and remove everything the engine now owns: the batch dimension, the attention implementation, the KV cache, the generation loop, sampling. What remains is the model’s layers, wired to vLLM’s shared building blocks and written against a flat batch. Then you register it. Seeing what a model must not do is the clearest summary of what the engine does, which makes this a fitting place to review the whole course.

A picture#

flowchart LR
  HF[":huggingface: <b>Reference model code</b><br/><small>batch × seq tensors, own attention,<br/>own cache, generate()</small>"] --> STRIP[":i-wrench: <b>Remove</b><br/><small>padding and masks, cache handling,<br/>generation loop, training code</small>"]
  STRIP --> ADAPT[":i-code: <b>Adapt</b><br/><small>flat input_ids + positions,<br/>vLLM Attention and linear layers,<br/>prefix on every module</small>"]
  ADAPT --> REG[":i-list-checks: <b>Register</b><br/><small>in-tree, or via a plugin</small>"]
  REG --> RUN[":vllm: <b>The engine does the rest</b><br/><small>scheduling, paging, batching,<br/>sampling, streaming</small>"]
  class HF neutral
  class STRIP warn
  class ADAPT compute
  class REG queue
  class RUN compute

How it really works#

What vLLM’s model interface looks like#

A vLLM model is a torch.nn.Module with a specific shape. The contributor guide lists the requirements; each one corresponds to something from an earlier lesson.

1. A uniform constructor, with a prefix on every module.

Python
class MyModelForCausalLM(nn.Module):
    def __init__(self, *, vllm_config: VllmConfig, prefix: str = ""):
        super().__init__()
        self.model = MyModel(vllm_config, prefix=f"{prefix}.model")

The guide gives two reasons for the prefix. Runtime support: “vLLM’s attention operators are registered in a model’s state by their full names. Each attention operator must have a unique prefix as its layer name to avoid conflicts” — this is how each layer finds its own KV cache tensor and how layers are grouped (More Than One Kind of Cache). Non-uniform quantization: a checkpoint may quantise some layers and not others, and the prefix is what a layer uses to look up its own treatment (Quantization in vLLM).

2. A forward function over a flat batch.

Python
def forward(
    self,
    input_ids: torch.Tensor | None,
    positions: torch.Tensor,
    intermediate_tensors: IntermediateTensors | None = None,
    inputs_embeds: torch.Tensor | None = None,
) -> torch.Tensor:

The guide: “treat input_ids and positions as flattened tensors with a single batch size dimension, without a max-sequence length dimension.” That is the flat batch from Building the Batch. There is no attention mask argument and no padding: request boundaries are the attention backend’s business.

intermediate_tensors exists for pipeline parallelism: a later stage receives the previous stage’s activations instead of token IDs. inputs_embeds is the door through which image and audio embeddings arrive (Multimodal Inputs).

3. An embed_input_ids method, so the engine can obtain text embeddings separately and overwrite placeholder positions before the layers run.

4. vLLM’s layers instead of your own. Attention becomes vllm.model_executor.layers.attention.Attention, which carries no mathematics of its own; it looks up the selected backend, the layer’s cache tensor and this step’s metadata from a forward context the model runner sets, and calls the backend (Attention Backends). Linear layers become vLLM’s parallel linear layers, which create only this GPU’s slice of each weight and know how to be quantised.

What you delete#

In the reference codeWhy it goesWho does it now
past_key_values threading through every layerThe cache is not a Python object passed aroundBlock pool + attention backend
Attention masks, padding, attention_maskThe batch is flat; nothing is paddedquery_start_loc, seq_lens
The scaled-dot-product attention itselfMust read a paged cacheThe attention backend
generate(), beam search, stopping criteriaOne step at a time, for many requestsScheduler + sampler
Logits processors, temperature, top-pBatched, in a fixed orderSampler
Training branches, dropout, lossInference only—
Loading and device placementSharded and quantised at constructionModel loader

The guide’s instruction for the forward method is simply: “remove any unnecessary code, such as training-specific code.”

Tensor parallelism comes from the layers#

You do not write tensor-parallel code. You use the layers that implement it:

LayerSplitsCommunication
ColumnParallelLinearOutput features across GPUsNone; outputs stay split
RowParallelLinearInput features across GPUsAn all-reduce to sum partial results
QKVParallelLinearThe fused query/key/value projection, by headNone
VocabParallelEmbeddingThe embedding tableAn all-reduce

An attention block is a column-parallel projection followed by a row-parallel one; an MLP is the same pair. Each pair costs one all-reduce, which is where “two collectives per layer” in Parallelism comes from. Because these layers build only their own shard, the model never exists in full on any one GPU (Executor, Worker, Model Runner).

Registering it#

In-tree, for a contribution upstream: add the class under vllm/model_executor/models/ and an entry in the model registry mapping the architecture name from config.json to a module and class.

Out-of-tree, for a private or experimental model: a general plugin (Plugins and Extension Points):

Python
def register():
    from vllm import ModelRegistry
    if "MyModelForCausalLM" not in ModelRegistry.get_supported_archs():
        ModelRegistry.register_model(
            "MyModelForCausalLM", "my_package.my_model:MyModelForCausalLM"
        )

Either way the key is the architectures entry of the model’s Hugging Face configuration. That string is how vLLM decides which class to instantiate for a given checkpoint.

Capabilities are interfaces#

Optional features are declared by inheriting marker interfaces, which the engine checks:

InterfaceDeclares
SupportsLoRALayers can take adapters
SupportsMultiModalThe model accepts image, audio or video and provides a processor
SupportsPPThe model can be split into pipeline stages
SupportsTranscriptionThe model can back the audio transcription API

A multimodal model additionally supplies the processor that computes placeholder counts and a method that turns media tensors into embeddings. The contributor documentation has a separate, longer guide for that; it is the bulk of the work for a vision-language model.

When the model is not a plain transformer#

If the architecture has layers that keep state differently — a sliding window, a state-space block, a new compressed attention — it needs more than a model class: a KV cache spec that tells the engine how much each layer stores, possibly a new single-type cache manager, and an attention backend that understands the layout. This is the path every recent hybrid model took into vLLM, and it is why a brand-new architecture often arrives first as “runs, but treats everything as full attention” and gets efficient a release or two later.

Testing it#

The guide’s testing page asks for, at minimum, a comparison against the reference implementation. A practical sequence:

  1. Greedy equivalence. Same prompt, temperature 0, compare token IDs with Hugging Face generate(). Small numerical differences are expected at 16-bit; a divergence in the first few tokens is a bug.
  2. Batch independence. One request alone versus the same request among thirty others. A difference means request boundaries are leaking.
  3. Chunked prefill. A prompt longer than max_num_batched_tokens. Wrong positions in the second chunk show up here.
  4. Prefix cache. The same prompt twice; the second must match the first.
  5. Tensor parallelism, if supported: -tp 2 against -tp 1.
  6. -O0 against -O2. A difference points at compilation or CUDA graph capture, not at your model.

Each test isolates one of the engine’s assumptions about a model. Failing one tells you which assumption your code violates.

Code#

The contract of a row-parallel linear layer: each rank holds a slice of the weight, computes a partial result, and the sum over ranks equals the full layer.

Go
package main

import "fmt"

// matVec computes W.x for W given as rows.
func matVec(w [][]float64, x []float64) []float64 {
	out := make([]float64, len(w))
	for i, row := range w {
		for j, v := range row {
			out[i] += v * x[j]
		}
	}
	return out
}

func main() {
	const (
		dOut = 3
		dIn  = 8
		tp   = 4
	)
	// The full weight, as it exists in the checkpoint: dOut x dIn.
	w := make([][]float64, dOut)
	for i := range w {
		w[i] = make([]float64, dIn)
		for j := range w[i] {
			w[i][j] = float64((i+1)*(j+2)%7) - 3
		}
	}
	x := []float64{1, -2, 0.5, 3, -1, 2, 0.25, -0.5}

	full := matVec(w, x)

	// Row-parallel: the INPUT dimension is split. Rank r is built with only its own
	// columns of W and receives only its own slice of x (the output of the preceding
	// column-parallel layer, which is already split the same way).
	shard := dIn / tp
	sum := make([]float64, dOut)
	fmt.Printf("full weight %dx%d = %d numbers; each of %d ranks holds %dx%d = %d\n\n",
		dOut, dIn, dOut*dIn, tp, dOut, shard, dOut*shard)
	for r := 0; r < tp; r++ {
		lo, hi := r*shard, (r+1)*shard
		wr := make([][]float64, dOut)
		for i := range wr {
			wr[i] = w[i][lo:hi] // only this slice is ever loaded on rank r
		}
		partial := matVec(wr, x[lo:hi])
		fmt.Printf("rank %d: columns %d-%d  partial = %v\n", r, lo, hi-1, partial)
		for i := range sum {
			sum[i] += partial[i] // the all-reduce
		}
	}

	fmt.Printf("\nall-reduce (sum over ranks): %v\n", sum)
	fmt.Printf("single-GPU result:           %v\n", full)
	fmt.Println("equal:", fmt.Sprint(sum) == fmt.Sprint(full))
}

No rank ever holds the whole matrix, and no rank needs to. That is the sharding-at-construction rule in twelve lines: the checkpoint’s full tensor is read once, and each process keeps only its columns.

Remember this#

  • A vLLM model is the reference model minus batching, attention, caching, generation and sampling.
  • Every module takes a prefix; it identifies the layer’s cache and its quantization.
  • forward takes flat input_ids and positions, with no padding and no mask.
  • Use vLLM’s Attention and parallel linear layers; tensor parallelism and sharded loading come with them.
  • The architectures string in the model’s config selects the class; register in-tree or through a plugin.
  • LoRA, multimodal and pipeline-parallel support are declared with marker interfaces.
  • Non-standard layers need a KV cache spec and often a new attention backend.
  • Test against the reference greedily, then in a batch, across chunks, with caching, with TP and at -O0.

Try it#

  1. In the program, change tp to 2 and to 8. How many numbers does each rank hold? What constraint does dIn have to satisfy?
  2. Open vllm/model_executor/models/opt.py, one of the smallest models in the tree, next to Hugging Face’s modeling_opt.py. List five things present in the second and absent from the first.
  3. Pick a model you use and find its architectures string in config.json. Then find the line in vLLM’s model registry that maps it to a class.

Check yourself#

  1. Why does a vLLM model’s forward have no attention-mask argument?
  2. What are the two uses of the prefix constructor argument?
  3. A new model gives correct output alone but wrong output in a batch. Which assumption is most likely violated?

Sources#

Checked on 5 October 2026 against main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom