The idea in one minute#
You rarely need to fork vLLM to change it. Almost every component you have read about in this course has a registered replacement mechanism: models, hardware platforms, attention backends, the scheduler class, logits processors, tool and reasoning parsers, LoRA resolvers, KV connectors, speculative proposers, metrics loggers and HTTP routes. They come in two styles. Entry-point plugins are Python packages that vLLM discovers and loads automatically in every one of its processes. Named classes are passed by module path on the command line. Knowing which extension point fits a need saves you from maintaining a patched build against a project that releases every two weeks.
A picture#
flowchart TB PKG[":python: <b>Your package</b><br/><small>pip install my-vllm-plugin</small>"] -->|"entry_points"| DISC[":i-search: <b>Discovery</b><br/><small>load_general_plugins()</small>"] DISC --> API[":vllm: API server"] DISC --> CORE[":vllm: Engine core"] DISC --> WK[":vllm: Each worker"] API --> R1["routes, tool/reasoning parsers,<br/>LoRA resolvers, stat loggers"] CORE --> R2["scheduler class,<br/>KV connector (scheduler side)"] WK --> R3["model classes, platform,<br/>logits processors, proposer,<br/>KV connector (worker side)"] class PKG neutral class DISC queue class API,CORE,WK compute class R1,R2,R3 memory
How it really works#
Why plugins must load in every process#
A vLLM server is several processes
(Processes and Wires), and with the spawn
start method a child does not inherit the parent’s Python state. A model class registered only
in the process you launched would be unknown to the worker that must instantiate it.
So plugin loading is something each process does for itself at startup. You saw one of the
call sites in EngineCore.__init__:
# plugins need to be loaded at the engine/scheduler level too
from vllm.plugins import load_general_plugins
load_general_plugins()This has one hard consequence for plugin authors, stated in the design document: the registered function must be re-entrant. “It can be called multiple times without causing issues. This is necessary because the function might be called multiple times in some processes.”
Discovery: Python entry points#
vLLM uses the standard entry_points mechanism. A package declares, in its metadata, a
function under a named group:
# setup.py
setup(
name="vllm_add_dummy_model",
packages=["vllm_add_dummy_model"],
entry_points={
"vllm.general_plugins": ["register_dummy_model = vllm_add_dummy_model:register"]
},
)
# vllm_add_dummy_model/__init__.py
def register():
from vllm import ModelRegistry
if "MyLlava" not in ModelRegistry.get_supported_archs():
ModelRegistry.register_model("MyLlava", "vllm_add_dummy_model.my_llava:MyLlava")Installing the package is enough; no flag is required. Three things to notice:
- The
if … not inguard is the re-entrancy requirement in practice. - The model is registered by a string path, not by importing the class. The import happens lazily in the process that needs it, which keeps startup fast and avoids importing GPU code in processes that have no GPU.
VLLM_PLUGINSfilters by plugin name: set it to a comma-separated list to load only those. Use it to bisect when an installed plugin is suspected of breaking startup.
The plugin groups#
| Group | Loaded in | Purpose |
|---|---|---|
vllm.general_plugins | Every process | Anything, usually registering out-of-tree models |
vllm.platform_plugins | Every process | A new hardware platform. The function returns the platform class’s path, or None if this machine is not that platform. |
vllm.io_processor_plugins | API server | Custom pre- and post-processing for pooling models |
vllm.stat_logger_plugins | API server | A custom metrics logger, a subclass of StatLoggerBase |
vllm.endpoint_plugins | API server only, and not by default | Extra HTTP routes on the server |
Platform plugins are how vLLM runs on hardware the core project does not ship support for. The README lists them: “Google TPUs, Intel Gaudi, IBM Spyre, Huawei Ascend, Rebellions NPU, Apple Silicon, MetaX GPU, and more.” A platform plugin supplies a worker, an attention backend, a device communicator and custom operations: essentially the bottom layer of Model Execution reimplemented for a new device, with the scheduler and KV cache manager reused unchanged.
Endpoint plugins need explicit opt-in because adding routes changes a server’s attack surface.
Named-class extension points#
These are selected by configuration rather than discovered:
| Extension | How to select | Interface | Lesson |
|---|---|---|---|
| Scheduler | --scheduler-cls my.module.MyScheduler | Subclass AsyncScheduler (subclassing plain Scheduler disables async scheduling) | Async Scheduling |
| Logits processor | --logits-processors my.module:MyProc | Batch-level processor with an update_state hook | Sampling and Structured Output |
| Tool parser | --tool-parser-plugin file.py + --tool-call-parser name | Extract calls from text and from a stream | Tool Calling and Reasoning |
| Reasoning parser | --reasoning-parser-plugin file.py + --reasoning-parser name | Split reasoning from content | same |
| LoRA resolver | Registered through a general plugin | Given an unknown model name, return an adapter or nothing | LoRA Adapters |
| KV connector | --kv-transfer-config '{"kv_connector": …, "kv_connector_module_path": …}' | KVConnectorBase_V1, both halves | Disaggregation and KV Connectors |
| Speculative proposer | --speculative-config '{"method": "custom_class", "model": "my.module.MyProposer"}' | A propose method | Speculative Decoding |
| Attention backend | --attention-backend CUSTOM with a registered class | The AttentionBackend interface | Attention Backends |
| Chat template | --chat-template file.jinja | Jinja2 | The Frontend |
How stable are these interfaces?#
Not equally, and the project says so in different ways.
- Covered by the deprecation policy: CLI flags, environment variables, the HTTP API, and the public Python API. These go through a staged pipeline tied to minor releases, with the removal version announced in advance.
- Explicitly unstable: the scheduler interface. Selecting a custom class logs: “This
scheduler interface is not public and compatibility may not be maintained.” The stat-logger
interface carries a similar warning in capital letters in
AsyncLLM’s docstring. - In between: model, connector and backend interfaces. They are used by out-of-tree code the project knows about, so breaking changes are announced, but they do change.
With a release every two weeks, a plugin that reaches into internals should pin a vLLM version and run its own tests against each new release before upgrading. A plugin that uses only registration functions and documented base classes travels much better.
Choosing the lightest tool#
| You want to… | Use | Not |
|---|---|---|
| Serve a model architecture vLLM lacks | A general plugin registering the model | A fork |
| Change how a known model formats tool calls | A tool-parser plugin | Editing the built-in parser |
| Ban or boost tokens per request | logit_bias, allowed-token lists, or a logits processor | Custom sampling code |
| Add an authentication header check | A reverse proxy in front | An endpoint plugin |
| Send metrics somewhere else | Scrape /metrics; a stat-logger plugin only if you need push | Patching the loggers |
| Load adapters on demand from your registry | A LoRA resolver | Runtime-loading endpoints on each replica |
| Prioritise some tenants | --scheduling-policy priority and a gateway that sets priority | A custom scheduler |
| Route requests by cached prefix | KV events and an external router | Anything inside vLLM |
The pattern in that table: prefer configuration to plugins, plugins to custom classes, and anything to a fork.
Code#
The two properties that make a plugin registry safe across processes: registration is re-entrant, and it stores a path to resolve later rather than the object itself.
package main
import (
"fmt"
"sort"
"strings"
)
// registry maps an architecture name to the PATH of the class that implements it.
// Nothing is imported until someone asks for it.
type registry struct {
paths map[string]string
loaded map[string]bool
log []string
}
func newRegistry() *registry {
return ®istry{paths: map[string]string{}, loaded: map[string]bool{}}
}
// register is re-entrant: calling it again with the same name is a no-op.
func (r *registry) register(arch, path string) {
if _, ok := r.paths[arch]; ok {
r.log = append(r.log, " skip "+arch+": already registered")
return
}
r.paths[arch] = path
r.log = append(r.log, " register "+arch+" -> "+path)
}
// resolve "imports" the class the first time it is needed in this process.
func (r *registry) resolve(arch string) string {
path, ok := r.paths[arch]
if !ok {
return "error: unknown architecture " + arch
}
if !r.loaded[arch] {
r.loaded[arch] = true
return "import " + path + " (first use in this process)"
}
return "already imported"
}
type plugin struct {
name string
fn func(*registry)
}
// loadPlugins is what every vLLM process runs at startup.
// allow mirrors the VLLM_PLUGINS environment variable ("" = load all).
func loadPlugins(r *registry, plugins []plugin, allow string) {
allowed := map[string]bool{}
for _, n := range strings.Split(allow, ",") {
if n != "" {
allowed[n] = true
}
}
sort.Slice(plugins, func(i, j int) bool { return plugins[i].name < plugins[j].name })
for _, p := range plugins {
if len(allowed) > 0 && !allowed[p.name] {
r.log = append(r.log, " filtered out "+p.name)
continue
}
p.fn(r)
}
}
func main() {
plugins := []plugin{
{"register_dummy_model", func(r *registry) { r.register("MyLlava", "my_pkg.my_llava:MyLlava") }},
{"register_other", func(r *registry) { r.register("OtherLM", "other_pkg.model:OtherLM") }},
}
// Each process has its own registry: nothing is shared across spawn().
for _, proc := range []string{"APIServer", "EngineCore", "Worker_TP0"} {
r := newRegistry()
fmt.Println(proc)
loadPlugins(r, plugins, "")
loadPlugins(r, plugins, "") // some processes load twice; this must be harmless
for _, l := range r.log {
fmt.Println(l)
}
if strings.HasPrefix(proc, "Worker") {
fmt.Println(" build model:", r.resolve("MyLlava"))
} else {
fmt.Println(" (this process never imports the model class)")
}
fmt.Println()
}
r := newRegistry()
fmt.Println("with VLLM_PLUGINS=register_other")
loadPlugins(r, plugins, "register_other")
for _, l := range r.log {
fmt.Println(l)
}
fmt.Println(" build model:", r.resolve("MyLlava"))
}Every process registers both plugins, twice, without harm; only the worker ever imports the model code. The last run shows the filter doing its job, and the error you would see if you filtered out the plugin your model needs.
Remember this#
- Plugins are loaded independently in every vLLM process, so registration must be re-entrant.
- Entry-point groups: general (models), platform (hardware), IO processor, stat logger, endpoint (off by default).
- Registration stores a path; the class is imported lazily where it is used.
VLLM_PLUGINSlimits which plugins load.- Scheduler, logits processors, parsers, connectors, proposers and attention backends are selected by name or path in configuration.
- Flags and the HTTP API follow a deprecation policy; the scheduler and stat-logger interfaces are explicitly unstable.
- Prefer configuration over plugins, plugins over custom classes, and all of them over a fork.
Try it#
- Remove the “already registered” guard from
registerso that it overwrites. Is the result still correct here? Describe a registration function for which double-calling would not be harmless. - In an environment with vLLM installed, run
python -c "from importlib.metadata import entry_points; print(entry_points(group='vllm.general_plugins'))". What is installed? - Write (on paper) the
registerfunction for a platform plugin that returnsNoneunless an environment variable is set. Why must it returnNonerather than raise?
Check yourself#
- Why must a plugin’s registration function be safe to call more than once?
- Why does
ModelRegistry.register_modeltake a string path instead of the class? - You need per-tenant priority. Which extension mechanism should you reach for first, and which last?
Sources#
Checked on 5 October 2026 against main at commit 0c16eee.
- Plugin system (design)
- Endpoint plugins
- LoRA resolver plugins
- IO processor plugins
- Deprecation policy
vllm/config/scheduler.py— the custom scheduler warning