The idea in one minute#
The engine core is a while loop. Each turn it empties its inbox, asks the scheduler what to
run, sends that to the workers, waits for sampled tokens, tells the scheduler what happened,
and posts the results. That is one step, and one step produces at most one token per
request. In the default configuration the loop does something subtler: it plans step N+1
while the GPU is still computing step N, so the GPU never waits for Python. Understanding
this loop is understanding where every millisecond of per-token latency goes.
A picture#
flowchart TB
A[":i-mail: <b>Drain input queue</b><br/><small>ADD, ABORT, UTILITY</small>"] --> B{"has work?"}
B -- "no" --> A
B -- "yes" --> C[":i-list-checks: <b>scheduler.schedule()</b><br/><small>pick tokens, allocate blocks</small>"]
C --> D[":i-split: <b>executor.execute_model()</b><br/><small>non-blocking, returns a future</small>"]
D --> E[":i-code: <b>get_grammar_bitmask()</b><br/><small>runs while the GPU works</small>"]
E --> F[":nvidia: <b>wait for forward pass</b><br/><small>then sample_tokens()</small>"]
F --> G[":i-ban: <b>process aborts</b><br/><small>that arrived meanwhile</small>"]
G --> H[":i-check: <b>scheduler.update_from_output()</b><br/><small>append tokens, stop checks, free</small>"]
H --> I[":i-message-square: <b>output queue</b><br/><small>to the output thread</small>"]
I --> A
class A,I io
class B,C,G,H queue
class D,E neutral
class F computeHow it really works#
The loop itself#
EngineCoreProc.run_busy_loop in vllm/v1/engine/core.py:
def run_busy_loop(self):
"""Core busy loop of the EngineCore."""
while self._handle_shutdown():
# 1) Poll the input queue until there is work to do.
self._process_input_queue()
# 2) Step the engine core and return the outputs.
self._process_engine_step()
raise SystemExit“Busy” has a precise meaning. _process_input_queue blocks on the queue only when there is
nothing to do:
def has_work(self) -> bool:
"""Returns true if the engine should be stepped."""
return (
self.engines_running
or self.scheduler.has_requests()
or bool(self.batch_queue)
)With no requests the engine sleeps on input_queue.get() and uses no CPU. The moment one
request exists, the loop spins without ever sleeping until every request is finished. An idle
vLLM server shows 0% CPU for the engine core; a loaded one shows 100% of one core. Both are
correct.
One step, the simple way#
EngineCore.step() is the version to learn first:
def step(self) -> tuple[dict[int, EngineCoreOutputs], bool]:
if not self.scheduler.has_requests():
return {}, False
scheduler_output = self.scheduler.schedule(self._should_throttle_prefills())
future = self.model_executor.execute_model(scheduler_output, non_block=True)
grammar_output = self.scheduler.get_grammar_bitmask(scheduler_output)
with (...):
model_output = future.result()
if model_output is None:
model_output = self.model_executor.sample_tokens(grammar_output)
# Before processing the model output, process any aborts that happened
# during the model execution.
self._process_aborts_queue()
engine_core_outputs = self.scheduler.update_from_output(
scheduler_output, model_output
)
return engine_core_outputs, scheduler_output.total_num_scheduled_tokens > 0Five things happen, in this order:
| # | Call | What it does | Runs on |
|---|---|---|---|
| 1 | scheduler.schedule() | Chooses how many tokens of which requests run; allocates KV blocks; returns a SchedulerOutput | CPU, engine core |
| 2 | execute_model(non_block=True) | Ships the SchedulerOutput to the workers and returns at once with a future | — |
| 3 | get_grammar_bitmask() | For requests using structured output, computes which tokens the grammar allows next | CPU, while the GPU runs |
| 4 | future.result() then sample_tokens() | Waits for the forward pass, then sends the bitmask and asks the workers to sample | GPU |
| 5 | update_from_output() | Appends each new token to its request, checks stop conditions, frees blocks of finished requests, builds EngineCoreOutputs | CPU, engine core |
Why the model call is split in two. The forward pass does not need to know which tokens a
JSON grammar permits; only sampling does. So execute_model starts the forward pass and
returns None, the engine core computes the grammar bitmask on the CPU during the forward
pass, and sample_tokens(grammar_output) delivers it just in time. Structured output
therefore costs CPU time that mostly overlaps with GPU time rather than adding to it.
Why aborts are processed between 4 and 5. A forward pass takes milliseconds to hundreds of
milliseconds. Any client that disconnected during it has an abort waiting. Applying those
aborts before update_from_output means the scheduler never appends a token to — or
allocates another block for — a request nobody is waiting for.
One step, the default way: the batch queue#
step() has a flaw. Between “forward pass finished” and “next forward pass started”, the GPU
is idle while Python runs update_from_output and then schedule. That is typically one to a
few milliseconds, on every token.
So the engine keeps a small queue of batches that have been scheduled but not finished:
@property
def max_concurrent_batches(self) -> int:
# PP requires PP-size concurrent batches to fill the pipeline.
# Async scheduling requires 2 concurrent batches to overlap.
pp_size = self.parallel_config.pipeline_parallel_size
if self.scheduler_config.async_scheduling:
if self.use_v2_model_runner:
return pp_size + 1
if pp_size <= 1:
return 2
return pp_sizeAsync scheduling is on by default (it is disabled automatically for embedding models and for a
few speculative-decoding methods). The queue size is therefore 2 on an ordinary single-GPU
server, and the engine uses step_with_batch_queue instead of step:
1. If the queue is not full and the scheduler has requests:
schedule a new batch, start it (non-blocking), push it on the queue.
If the queue still has room, return immediately with no output.
2. Otherwise block on the OLDEST batch in the queue.
3. update_from_output() for that batch; return its outputs.The source comment states the priority: “fulfilling the batch queue has a higher priority than getting model outputs.”
The effect is that batch N+1 is scheduled and sent to the worker while batch N is still on
the GPU. When N finishes, N+1 is already waiting in the worker and starts with no gap.
synchronous GPU: [ step 1 ]......[ step 2 ]......[ step 3 ]
CPU: [u+s] [u+s]
async scheduling GPU: [ step 1 ][ step 2 ][ step 3 ][ step 4 ]
CPU: [s2] [u1+s3] [u2+s4] [u3+s5]
u = update_from_output, s = scheduleThere is a catch, and it shapes the scheduler. When batch N+1 is being scheduled, the tokens
that batch N will sample do not exist yet. The scheduler has to schedule the next decode
step for a request without knowing its previous output token. It does this with
placeholders; the worker, which does know the token, fills it in. That mechanism is the
subject of Async Scheduling.
Pipeline parallelism uses the same queue for a different reason. With the model split across
two GPUs in series, GPU 0 is free as soon as it has handed its activations to GPU 1. A queue of
size PP keeps every stage busy with a different batch, removing what the code calls “pipeline
bubbles”.
What happens after a step#
def _process_engine_step(self) -> bool:
outputs, model_executed = self.step_fn()
for output in outputs.items() if outputs else ():
self.output_queue.put_nowait(output)
self.post_step(model_executed)
# If no model execution happened but there is still scheduler work
# (e.g. WAITING_FOR_REMOTE_KVS or delayed KV connector frees), yield
# the GIL briefly to allow background transfer threads to make progress.
if not model_executed and self.scheduler.has_requests():
time.sleep(0.001)
return model_executed- Outputs are keyed by client index: a step’s results are split per API server, so each API server receives only its own requests.
post_stepcollects draft tokens for speculative decoding when async scheduling is off.- The one-millisecond sleep is the single place the busy loop yields. It applies when requests exist but none could run — for instance when every waiting request is waiting for KV data to arrive from another machine. Without it the loop would spin at 100% CPU and starve the threads doing that transfer.
Messages other than requests#
_handle_client_request dispatches on the type byte from
Processes and Wires:
| Type | Action |
|---|---|
ADD | scheduler.add_request(request) — the request joins the waiting queue |
ABORT | scheduler.finish_requests(ids, FINISHED_ABORTED) — idempotent, so the duplicate on the fast aborts queue is harmless |
UTILITY | Look up a method on the engine core by name, call it, post the result |
EXECUTOR_FAILED | raise RuntimeError("Executor failed.") — a worker died; take the engine down |
WAKEUP | Nothing. It exists only to unblock input_queue.get() |
Utility calls run on the main loop thread, between steps. That is why operations such as loading a LoRA adapter or resetting the prefix cache are safe without locks: no step is in progress while they run. It is also why a slow utility call stalls generation for everyone.
Pause, sleep and wake#
Three utility methods change the loop’s behaviour without stopping the process.
pause_scheduler(mode) has three modes:
| Mode | In-flight requests | New requests |
|---|---|---|
abort (default) | Aborted immediately | Queued until resume |
wait | Allowed to finish | Queued until resume |
keep | Frozen in place; no steps run | Queued until resume |
sleep(level) builds on it, to release GPU memory while keeping the process alive:
| Level | Effect | Use |
|---|---|---|
| 0 | Pause scheduling only. No memory changes. | A brief hold |
| 1 | Move weights to CPU RAM, discard the KV cache | Share a GPU between models that take turns; waking is fast |
| 2 | Discard all GPU memory | Longest idle; waking reloads weights from disk |
wake_up() reverses it and resumes scheduling once all memory is back. The server exposes
these as /sleep, /wake_up and /is_sleeping when started with --enable-sleep-mode.
Reinforcement-learning loops use pause with wait to swap in new weights between batches of
generation.
Start-up of the engine core#
Before the loop runs, EngineCore.__init__ does the work whose results you read in the
startup log:
- Create the executor, which starts the workers and loads the model.
_initialize_kv_caches: ask the workers for each layer’s cache requirements, profile free memory, compute the number of blocks, tell the workers to allocate them, then compile and warm up (Sizing the Cache).- Create the structured output manager and the scheduler.
- Choose
steporstep_with_batch_queue. freeze_gc_heap(): mark everything allocated so far as permanent so the garbage collector never scans it again.
Two decisions are taken here that surprise people later: if the model has no KV cache (an encoder-only embedding model), chunked prefill is switched off; if any layer attends non-causally, both chunked prefill and prefix caching are switched off, because both assume a token’s KV depends only on the tokens before it.
Code#
How much does overlapping scheduling with the GPU buy? This program times both loops for a range of batch sizes. CPU cost per step grows with the number of requests; GPU cost grows more slowly, so the idle gap matters most exactly when the server is busy.
package main
import "fmt"
func main() {
fmt.Printf("%8s %10s %10s %14s %14s %9s\n",
"requests", "cpu ms", "gpu ms", "sync tok/s", "async tok/s", "gain")
for _, n := range []int{1, 8, 32, 128, 256, 512} {
// Rough shapes, not measurements: Python work is linear in batch size,
// a decode forward pass has a fixed cost plus a small per-request cost.
cpu := 0.25 + 0.012*float64(n) // schedule + update_from_output
gpu := 8.0 + 0.03*float64(n) // one decode step
syncStep := cpu + gpu // the GPU waits while Python plans the next step
asyncStep := gpu // planning overlaps the previous forward pass
if cpu > gpu {
asyncStep = cpu // Python is now the bottleneck
}
syncTok := float64(n) / syncStep * 1000
asyncTok := float64(n) / asyncStep * 1000
fmt.Printf("%8d %10.2f %10.2f %14.0f %14.0f %8.1f%%\n",
n, cpu, gpu, syncTok, asyncTok, 100*(asyncTok/syncTok-1))
}
}The gain grows with concurrency, which is why the feature became the default. Notice the
if cpu > gpu branch as well: once Python’s per-step work exceeds the forward pass, more GPU
does not help. That happens with very small models at very high concurrency, and it is the
motivation for the Rust frontend and for moving work out of the engine core’s main thread.
Remember this#
- The engine core is a loop: drain inbox, schedule, execute, update, post.
- It blocks only when it has no requests; otherwise it uses one full core.
- The model call is split so grammar bitmasks are computed while the GPU runs.
- Aborts are applied right after each forward pass, before outputs are processed.
- By default a batch queue of size 2 lets step
N+1be scheduled while stepNis on the GPU. - Utility calls (LoRA, cache reset, sleep) run between steps on the loop thread.
sleeplevels 0, 1, 2 trade wake-up time for released GPU memory.
Try it#
- In the program, change the GPU cost to
2.0 + 0.01*n(a much smaller model). At which batch size does Python become the bottleneck? - On a running server, watch the engine core’s CPU with
topwhile idle and while serving one long request. Explain both readings usinghas_work(). - Start a server with
--enable-sleep-mode, POST to/sleep?level=1, and watchnvidia-smi. Then POST to/wake_up. How long does each take, and where did the weights go in between?
Check yourself#
- In
step(), why isget_grammar_bitmaskcalled afterexecute_modeland not before it? - What information is missing when the engine schedules batch
N+1before batchNhas finished? - Why is it safe to add a LoRA adapter without a lock?
Sources#
Checked on 5 October 2026 against main at commit 0c16eee.
vllm/v1/engine/core.py—run_busy_loop,step,step_with_batch_queue,sleepvllm/config/vllm.py—max_concurrent_batches, when async scheduling is enabled- Sleep mode