Pidoku

Multimodal Inputs

Advanced 45 min Difficulty 4/5 Lesson 03 of 05

Prerequisites The Frontend, Building the Batch, Prefix Caching

The idea in one minute#

A vision-language model does not read pixels. An encoder turns an image into a few hundred or a few thousand vectors that look, to the language model, like token embeddings. vLLM makes this fit its text machinery with one device: in the token list, an image is a run of placeholder tokens, one per vector. Everything built for text — the token budget, chunked prefill, block hashes, the prefix cache — then works on the placeholders unchanged. Around that device sit three pieces of plumbing: fetching and preprocessing media on the CPU, running the encoder on the GPU under its own budget, and three caches so the same image is never processed twice.

A picture#

flowchart LR
  URL[":i-globe: <b>image_url</b><br/><small>https, data: or file:</small>"] --> F[":i-cpu: <b>Fetch + decode</b><br/><small>API server, thread pool</small>"]
  F --> PR[":huggingface: <b>Processor</b><br/><small>resize, tile, to pixel tensor</small>"]
  PR --> PH["prompt: ... &lt;img&gt;×576 ...<br/><small>576 placeholder tokens</small>"]
  PR --> PC[(":i-database: <b>Processor cache</b><br/><small>by content hash</small>")]
  PH --> SCH[":i-list-checks: <b>Scheduler</b><br/><small>placeholders are ordinary tokens;<br/>encoder has its own budget</small>"]
  SCH --> ENC[":nvidia: <b>Vision encoder</b><br/><small>runs when the chunk reaches the image</small>"]
  ENC --> EC[(":i-layers: <b>Encoder cache</b><br/><small>embeddings, until consumed</small>")]
  EC --> LM[":nvidia: <b>Language model</b><br/><small>placeholder embeddings overwritten</small>"]
  class URL neutral
  class F,PR compute
  class PH,SCH queue
  class PC,EC memory
  class ENC,LM compute

How it really works#

The request#

The OpenAI chat format carries media as content parts:

JSON
{"role": "user", "content": [
  {"type": "text", "text": "What is in this image?"},
  {"type": "image_url", "image_url": {"url": "https://example.com/cat.jpg"}}
]}

video_url, input_audio and audio_url work the same way. A URL may be https:, a base64 data: URL, or — only if you enable it — a local file: path.

Step 1: fetch and decode, in the API server#

The API server downloads and decodes media on a thread pool (VLLM_MEDIA_LOADING_THREAD_COUNT, default 8). This is CPU work and network I/O that never touches the engine core.

warning

A server that fetches URLs on behalf of clients can be made to fetch internal URLs: cloud metadata endpoints, admin panels, other services on the cluster network. This is server-side request forgery. The documentation’s first note on this page is to set --allowed-media-domains. Local file access is off unless you pass --allowed-local-media-path, and should point at a directory containing nothing else.

Limits are part of the model’s contract with clients: --limit-mm-per-prompt.image 2 rejects requests with more than two images. Because every image can add thousands of tokens, these limits also bound memory.

Step 2: the processor and the placeholders#

Each model family has a multimodal processor that does two things:

  1. Turns the decoded image into the tensor the encoder expects (resizing, splitting into tiles, normalising).
  2. Works out how many vectors the encoder will produce and replaces the single <image> marker in the prompt with that many placeholder tokens.
before:  <s>[INST]What is in this image?\n[IMG][/INST]
after:   [1, 3, 7493, 1681, 1294, 1593, 3937, 9551, <P>, <P>, ..., <P>, 4]

The count depends on the model and often on the image’s size and aspect ratio. It is the real cost of an image: each placeholder occupies a position in the context window and a slot in the KV cache in every layer, exactly like a word.

The processor also records where each item sits: an mm_position (offset and length in the token list) and an identifier, a hash of the item’s content. Those two facts are all the engine needs from here on.

Step 3: three caches#

Processing can be slow; the design document notes that some processors “are very slow”. vLLM caches at three levels.

CacheWhereKeyed byHoldsSaves
Processor cacheAPI server and engine coreContent hashThe processed pixel tensorRe-running the processor
Encoder cacheWorker (GPU)Content hashThe encoder’s output embeddingsRe-running the encoder within a request’s lifetime
Prefix cacheEngine core + GPUBlock hash including the content hashThe language model’s KV for the placeholder blocksRe-running the language model over the image

The processor cache defaults to 4 GiB (--mm-processor-cache-gb). Its interesting property is how it avoids resending data. The API server keeps a “shadow” cache of keys; the engine core keeps the values. When an image the engine already holds is sent again, only its hash crosses the socket. If the two fall out of step, the engine reports a retryable multimodal_cache_miss, the API server forgets the stale key, and the client’s retry carries the data.

The prefix cache needed one addition. Placeholder tokens are identical for every image, so two different pictures of the same size would hash identically. Hence the extra key from Prefix Caching:

Python
# The block contains the current mm input. Include its offset
# relative to the start of the block so prefix-cache keys stay
# distinct when the same MM item appears at different positions
# within otherwise-identical placeholder blocks.
extra_keys.append(("mm", mm_feature.identifier, offset - start_token_idx))

With that, a second question about the same image reuses the image’s KV blocks and computes only the new text.

Step 4: scheduling the encoder#

The scheduler treats placeholders as tokens but must also decide when the encoder runs. A vision encoder is a sizeable model; running ten of them in one step would blow the step time. So the encoder has its own budget and its own cache accounting:

Python
# Encoder-related.
scheduled_encoder_inputs: dict[str, list[int]] = {}
encoder_compute_budget = self.max_num_encoder_input_tokens

max_num_encoder_input_tokens and the encoder cache size both default to max_num_batched_tokens. For each request with media, _try_schedule_encoder_inputs checks whether this step’s token range overlaps an item that has not been encoded yet, and if so whether the encoder budget and cache can take it. If not, the request’s chunk is cut short to end just before the item, and the image waits for a later step.

One further rule: by default a chunk may end in the middle of an image’s placeholders. For models that need a whole image in one forward pass, --disable-chunked-mm-input forbids that: the config’s example is a request scheduled as TTTT in one step and IIIIIIIIII in the next rather than TTTTIIIII and IIIII.

Step 5: in the model runner#

When the batch reaches the worker (Building the Batch):

  1. For every item scheduled this step and not in the encoder cache, run the encoder and store the embeddings.
  2. Look up the ordinary embeddings for all input_ids, placeholders included.
  3. Overwrite the rows at placeholder positions with the corresponding slice of encoder output. With chunked prefill this may be only part of an image.
  4. Run the language model as usual.

Embeddings are evicted from the encoder cache once the request has consumed all of an item’s placeholders; the scheduler tells the worker which hashes to free (free_encoder_mm_hashes).

Making preprocessing cheaper#

Two optimisations on the data path are worth knowing because they change sizing.

Pixels travel as bytes. Older pipelines normalised on the CPU and shipped 16-bit floats. The current path keeps uint8 all the way to the device and applies normalisation there as one fused operation:

Before: CPU decode → CPU resize → CPU rescale (÷255) → CPU normalize
        → cast to bf16 → PCIe (2 B/elem) → GPU
After:  CPU decode → CPU resize → uint8 pixel_values
        → PCIe (1 B/elem) → GPU fused affine (fp32 compute) → bf16

Half the bytes over the socket and the bus, and less CPU.

The encoder can live elsewhere. A disaggregated-encoder deployment runs vision encoders in separate vLLM instances that publish embeddings through an encoder-cache connector; the language-model instances fetch them instead of encoding. It is the same idea as prefill/decode disaggregation (Disaggregation and KV Connectors), applied to a different stage.

Sizing#

  • CPU. Decoding and resizing are per-request CPU work in the API server. Image-heavy traffic is often CPU-bound before it is GPU-bound. Add API server processes (--api-server-count) and cores.
  • Context. max_model_len must cover text plus placeholders. The validation error says so: “Make sure that max_model_len is no smaller than the number of text tokens plus multimodal tokens.”
  • Memory. The encoder’s weights and its peak activations are part of non-KV memory (Sizing the Cache); a little memory is also reserved for moving tensors between processes.
  • Startup. Encoders have their own compilation and CUDA graph capture, reported separately in the init engine … took line.

Code#

What an image costs in tokens, context and cache. The formula is the common one for patch-based encoders: one vector per patch, optionally merged in small groups.

Go
package main

import "fmt"

type encoder struct {
	name  string
	patch int // pixels per patch side
	merge int // patches merged per side into one token (1 = none)
	tile  int // images larger than this are split into tiles (0 = resize instead)
}

func ceilDiv(a, b int) int { return (a + b - 1) / b }

// tokens returns the number of placeholder tokens for a w x h image.
func (e encoder) tokens(w, h int) int {
	if e.tile == 0 {
		// Fixed-resolution encoder: every image is resized to one size.
		side := 336 / e.patch
		return side * side / (e.merge * e.merge)
	}
	// Tiling encoder: a global thumbnail plus one grid of patches per tile.
	tiles := ceilDiv(w, e.tile) * ceilDiv(h, e.tile)
	perTile := (e.tile / e.patch) * (e.tile / e.patch) / (e.merge * e.merge)
	return (tiles + 1) * perTile
}

func main() {
	encoders := []encoder{
		{"fixed 336px, patch 14", 14, 1, 0},
		{"tiled 448px, patch 14, 2x2 merge", 14, 2, 448},
	}
	images := [][2]int{{336, 336}, {1024, 768}, {1920, 1080}, {3840, 2160}}

	const kvBytesPerToken = 131072 // an 8B-class language model, 16-bit cache
	const budget = 8192            // max_num_batched_tokens

	for _, e := range encoders {
		fmt.Println(e.name)
		fmt.Printf("  %-12s %8s %14s %18s\n", "image", "tokens", "KV cache", "steps to prefill")
		for _, im := range images {
			n := e.tokens(im[0], im[1])
			fmt.Printf("  %4dx%-7d %8d %10.0f MiB %18d\n",
				im[0], im[1], n, float64(n*kvBytesPerToken)/(1<<20), ceilDiv(n, budget))
		}
		fmt.Println()
	}

	n := encoders[1].tokens(1920, 1080)
	fmt.Printf("a 1920x1080 image costs %d tokens: about %d words of English in context,\n", n, n*3/4)
	fmt.Printf("and ten such images in one request need max_model_len of at least %d\n", 10*n)
}

The fixed-resolution encoder charges the same for a thumbnail and a 4K frame because it shrinks everything. The tiling encoder preserves detail and its cost grows with the pixel count. Knowing which kind your model has is the first step in setting --limit-mm-per-prompt and --max-model-len.

Remember this#

  • An image becomes a run of placeholder tokens, one per encoder output vector.
  • Placeholders are scheduled, chunked, counted and cached exactly like text tokens.
  • Media is fetched and preprocessed in the API server on the CPU; restrict the domains it may fetch from.
  • Three caches: processed pixels (by hash, default 4 GiB), encoder outputs (per request), and the KV prefix cache.
  • The content hash is part of the block hash, so the same image hits the prefix cache and different images never collide.
  • The encoder runs under its own per-step budget; a chunk is cut short if the encoder cannot take the next item.
  • Image traffic is often CPU-bound in the API server before it is GPU-bound.

Try it#

  1. Add an encoder with a 3×3 merge. How much does it save on a 4K image, and what do you expect it to cost in fine detail?
  2. Using the second encoder, what --limit-mm-per-prompt.image keeps a request within a 32,768-token context if images are at most 1920×1080 and text is at most 2,000 tokens?
  3. On a real multimodal server, send the same image twice with different questions and compare usage.prompt_tokens_details.cached_tokens. Then send a different image of the same size.

Check yourself#

  1. Why does representing an image as placeholder tokens let the scheduler treat it like text?
  2. Two different images produce identical placeholder tokens. What keeps the prefix cache from confusing them?
  3. Which process downloads an image URL, and what is the security risk of letting it fetch any URL?

Sources#

Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom