Pidoku

Plugins and Extension Points

Expert 35 min Difficulty 3/5 Lesson 01 of 03

Prerequisites Processes and Wires, Sampling and Structured Output

The idea in one minute#

You rarely need to fork vLLM to change it. Almost every component you have read about in this course has a registered replacement mechanism: models, hardware platforms, attention backends, the scheduler class, logits processors, tool and reasoning parsers, LoRA resolvers, KV connectors, speculative proposers, metrics loggers and HTTP routes. They come in two styles. Entry-point plugins are Python packages that vLLM discovers and loads automatically in every one of its processes. Named classes are passed by module path on the command line. Knowing which extension point fits a need saves you from maintaining a patched build against a project that releases every two weeks.

A picture#

flowchart TB
  PKG[":python: <b>Your package</b><br/><small>pip install my-vllm-plugin</small>"] -->|"entry_points"| DISC[":i-search: <b>Discovery</b><br/><small>load_general_plugins()</small>"]
  DISC --> API[":vllm: API server"]
  DISC --> CORE[":vllm: Engine core"]
  DISC --> WK[":vllm: Each worker"]
  API --> R1["routes, tool/reasoning parsers,<br/>LoRA resolvers, stat loggers"]
  CORE --> R2["scheduler class,<br/>KV connector (scheduler side)"]
  WK --> R3["model classes, platform,<br/>logits processors, proposer,<br/>KV connector (worker side)"]
  class PKG neutral
  class DISC queue
  class API,CORE,WK compute
  class R1,R2,R3 memory

How it really works#

Why plugins must load in every process#

A vLLM server is several processes (Processes and Wires), and with the spawn start method a child does not inherit the parent’s Python state. A model class registered only in the process you launched would be unknown to the worker that must instantiate it.

So plugin loading is something each process does for itself at startup. You saw one of the call sites in EngineCore.__init__:

Python
# plugins need to be loaded at the engine/scheduler level too
from vllm.plugins import load_general_plugins
load_general_plugins()

This has one hard consequence for plugin authors, stated in the design document: the registered function must be re-entrant. “It can be called multiple times without causing issues. This is necessary because the function might be called multiple times in some processes.”

Discovery: Python entry points#

vLLM uses the standard entry_points mechanism. A package declares, in its metadata, a function under a named group:

Python
# setup.py
setup(
    name="vllm_add_dummy_model",
    packages=["vllm_add_dummy_model"],
    entry_points={
        "vllm.general_plugins": ["register_dummy_model = vllm_add_dummy_model:register"]
    },
)

# vllm_add_dummy_model/__init__.py
def register():
    from vllm import ModelRegistry
    if "MyLlava" not in ModelRegistry.get_supported_archs():
        ModelRegistry.register_model("MyLlava", "vllm_add_dummy_model.my_llava:MyLlava")

Installing the package is enough; no flag is required. Three things to notice:

  • The if … not in guard is the re-entrancy requirement in practice.
  • The model is registered by a string path, not by importing the class. The import happens lazily in the process that needs it, which keeps startup fast and avoids importing GPU code in processes that have no GPU.
  • VLLM_PLUGINS filters by plugin name: set it to a comma-separated list to load only those. Use it to bisect when an installed plugin is suspected of breaking startup.

The plugin groups#

GroupLoaded inPurpose
vllm.general_pluginsEvery processAnything, usually registering out-of-tree models
vllm.platform_pluginsEvery processA new hardware platform. The function returns the platform class’s path, or None if this machine is not that platform.
vllm.io_processor_pluginsAPI serverCustom pre- and post-processing for pooling models
vllm.stat_logger_pluginsAPI serverA custom metrics logger, a subclass of StatLoggerBase
vllm.endpoint_pluginsAPI server only, and not by defaultExtra HTTP routes on the server

Platform plugins are how vLLM runs on hardware the core project does not ship support for. The README lists them: “Google TPUs, Intel Gaudi, IBM Spyre, Huawei Ascend, Rebellions NPU, Apple Silicon, MetaX GPU, and more.” A platform plugin supplies a worker, an attention backend, a device communicator and custom operations: essentially the bottom layer of Model Execution reimplemented for a new device, with the scheduler and KV cache manager reused unchanged.

Endpoint plugins need explicit opt-in because adding routes changes a server’s attack surface.

Named-class extension points#

These are selected by configuration rather than discovered:

ExtensionHow to selectInterfaceLesson
Scheduler--scheduler-cls my.module.MySchedulerSubclass AsyncScheduler (subclassing plain Scheduler disables async scheduling)Async Scheduling
Logits processor--logits-processors my.module:MyProcBatch-level processor with an update_state hookSampling and Structured Output
Tool parser--tool-parser-plugin file.py + --tool-call-parser nameExtract calls from text and from a streamTool Calling and Reasoning
Reasoning parser--reasoning-parser-plugin file.py + --reasoning-parser nameSplit reasoning from contentsame
LoRA resolverRegistered through a general pluginGiven an unknown model name, return an adapter or nothingLoRA Adapters
KV connector--kv-transfer-config '{"kv_connector": …, "kv_connector_module_path": …}'KVConnectorBase_V1, both halvesDisaggregation and KV Connectors
Speculative proposer--speculative-config '{"method": "custom_class", "model": "my.module.MyProposer"}'A propose methodSpeculative Decoding
Attention backend--attention-backend CUSTOM with a registered classThe AttentionBackend interfaceAttention Backends
Chat template--chat-template file.jinjaJinja2The Frontend

How stable are these interfaces?#

Not equally, and the project says so in different ways.

  • Covered by the deprecation policy: CLI flags, environment variables, the HTTP API, and the public Python API. These go through a staged pipeline tied to minor releases, with the removal version announced in advance.
  • Explicitly unstable: the scheduler interface. Selecting a custom class logs: “This scheduler interface is not public and compatibility may not be maintained.” The stat-logger interface carries a similar warning in capital letters in AsyncLLM’s docstring.
  • In between: model, connector and backend interfaces. They are used by out-of-tree code the project knows about, so breaking changes are announced, but they do change.

With a release every two weeks, a plugin that reaches into internals should pin a vLLM version and run its own tests against each new release before upgrading. A plugin that uses only registration functions and documented base classes travels much better.

Choosing the lightest tool#

You want to…UseNot
Serve a model architecture vLLM lacksA general plugin registering the modelA fork
Change how a known model formats tool callsA tool-parser pluginEditing the built-in parser
Ban or boost tokens per requestlogit_bias, allowed-token lists, or a logits processorCustom sampling code
Add an authentication header checkA reverse proxy in frontAn endpoint plugin
Send metrics somewhere elseScrape /metrics; a stat-logger plugin only if you need pushPatching the loggers
Load adapters on demand from your registryA LoRA resolverRuntime-loading endpoints on each replica
Prioritise some tenants--scheduling-policy priority and a gateway that sets priorityA custom scheduler
Route requests by cached prefixKV events and an external routerAnything inside vLLM

The pattern in that table: prefer configuration to plugins, plugins to custom classes, and anything to a fork.

Code#

The two properties that make a plugin registry safe across processes: registration is re-entrant, and it stores a path to resolve later rather than the object itself.

Go
package main

import (
	"fmt"
	"sort"
	"strings"
)

// registry maps an architecture name to the PATH of the class that implements it.
// Nothing is imported until someone asks for it.
type registry struct {
	paths  map[string]string
	loaded map[string]bool
	log    []string
}

func newRegistry() *registry {
	return &registry{paths: map[string]string{}, loaded: map[string]bool{}}
}

// register is re-entrant: calling it again with the same name is a no-op.
func (r *registry) register(arch, path string) {
	if _, ok := r.paths[arch]; ok {
		r.log = append(r.log, "  skip "+arch+": already registered")
		return
	}
	r.paths[arch] = path
	r.log = append(r.log, "  register "+arch+" -> "+path)
}

// resolve "imports" the class the first time it is needed in this process.
func (r *registry) resolve(arch string) string {
	path, ok := r.paths[arch]
	if !ok {
		return "error: unknown architecture " + arch
	}
	if !r.loaded[arch] {
		r.loaded[arch] = true
		return "import " + path + " (first use in this process)"
	}
	return "already imported"
}

type plugin struct {
	name string
	fn   func(*registry)
}

// loadPlugins is what every vLLM process runs at startup.
// allow mirrors the VLLM_PLUGINS environment variable ("" = load all).
func loadPlugins(r *registry, plugins []plugin, allow string) {
	allowed := map[string]bool{}
	for _, n := range strings.Split(allow, ",") {
		if n != "" {
			allowed[n] = true
		}
	}
	sort.Slice(plugins, func(i, j int) bool { return plugins[i].name < plugins[j].name })
	for _, p := range plugins {
		if len(allowed) > 0 && !allowed[p.name] {
			r.log = append(r.log, "  filtered out "+p.name)
			continue
		}
		p.fn(r)
	}
}

func main() {
	plugins := []plugin{
		{"register_dummy_model", func(r *registry) { r.register("MyLlava", "my_pkg.my_llava:MyLlava") }},
		{"register_other", func(r *registry) { r.register("OtherLM", "other_pkg.model:OtherLM") }},
	}

	// Each process has its own registry: nothing is shared across spawn().
	for _, proc := range []string{"APIServer", "EngineCore", "Worker_TP0"} {
		r := newRegistry()
		fmt.Println(proc)
		loadPlugins(r, plugins, "")
		loadPlugins(r, plugins, "") // some processes load twice; this must be harmless
		for _, l := range r.log {
			fmt.Println(l)
		}
		if strings.HasPrefix(proc, "Worker") {
			fmt.Println("  build model:", r.resolve("MyLlava"))
		} else {
			fmt.Println("  (this process never imports the model class)")
		}
		fmt.Println()
	}

	r := newRegistry()
	fmt.Println("with VLLM_PLUGINS=register_other")
	loadPlugins(r, plugins, "register_other")
	for _, l := range r.log {
		fmt.Println(l)
	}
	fmt.Println("  build model:", r.resolve("MyLlava"))
}

Every process registers both plugins, twice, without harm; only the worker ever imports the model code. The last run shows the filter doing its job, and the error you would see if you filtered out the plugin your model needs.

Remember this#

  • Plugins are loaded independently in every vLLM process, so registration must be re-entrant.
  • Entry-point groups: general (models), platform (hardware), IO processor, stat logger, endpoint (off by default).
  • Registration stores a path; the class is imported lazily where it is used.
  • VLLM_PLUGINS limits which plugins load.
  • Scheduler, logits processors, parsers, connectors, proposers and attention backends are selected by name or path in configuration.
  • Flags and the HTTP API follow a deprecation policy; the scheduler and stat-logger interfaces are explicitly unstable.
  • Prefer configuration over plugins, plugins over custom classes, and all of them over a fork.

Try it#

  1. Remove the “already registered” guard from register so that it overwrites. Is the result still correct here? Describe a registration function for which double-calling would not be harmless.
  2. In an environment with vLLM installed, run python -c "from importlib.metadata import entry_points; print(entry_points(group='vllm.general_plugins'))". What is installed?
  3. Write (on paper) the register function for a platform plugin that returns None unless an environment variable is set. Why must it return None rather than raise?

Check yourself#

  1. Why must a plugin’s registration function be safe to call more than once?
  2. Why does ModelRegistry.register_model take a string path instead of the class?
  3. You need per-tenant priority. Which extension mechanism should you reach for first, and which last?

Sources#

Checked on 5 October 2026 against main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom