Pidoku

The Engine Core Loop

Basic 50 min Difficulty 3/5 Lesson 04 of 04

Prerequisites AsyncLLM and the Way Back

The idea in one minute#

The engine core is a while loop. Each turn it empties its inbox, asks the scheduler what to run, sends that to the workers, waits for sampled tokens, tells the scheduler what happened, and posts the results. That is one step, and one step produces at most one token per request. In the default configuration the loop does something subtler: it plans step N+1 while the GPU is still computing step N, so the GPU never waits for Python. Understanding this loop is understanding where every millisecond of per-token latency goes.

A picture#

flowchart TB
  A[":i-mail: <b>Drain input queue</b><br/><small>ADD, ABORT, UTILITY</small>"] --> B{"has work?"}
  B -- "no" --> A
  B -- "yes" --> C[":i-list-checks: <b>scheduler.schedule()</b><br/><small>pick tokens, allocate blocks</small>"]
  C --> D[":i-split: <b>executor.execute_model()</b><br/><small>non-blocking, returns a future</small>"]
  D --> E[":i-code: <b>get_grammar_bitmask()</b><br/><small>runs while the GPU works</small>"]
  E --> F[":nvidia: <b>wait for forward pass</b><br/><small>then sample_tokens()</small>"]
  F --> G[":i-ban: <b>process aborts</b><br/><small>that arrived meanwhile</small>"]
  G --> H[":i-check: <b>scheduler.update_from_output()</b><br/><small>append tokens, stop checks, free</small>"]
  H --> I[":i-message-square: <b>output queue</b><br/><small>to the output thread</small>"]
  I --> A
  class A,I io
  class B,C,G,H queue
  class D,E neutral
  class F compute

How it really works#

The loop itself#

EngineCoreProc.run_busy_loop in vllm/v1/engine/core.py:

Python
def run_busy_loop(self):
    """Core busy loop of the EngineCore."""
    while self._handle_shutdown():
        # 1) Poll the input queue until there is work to do.
        self._process_input_queue()
        # 2) Step the engine core and return the outputs.
        self._process_engine_step()
    raise SystemExit

“Busy” has a precise meaning. _process_input_queue blocks on the queue only when there is nothing to do:

Python
def has_work(self) -> bool:
    """Returns true if the engine should be stepped."""
    return (
        self.engines_running
        or self.scheduler.has_requests()
        or bool(self.batch_queue)
    )

With no requests the engine sleeps on input_queue.get() and uses no CPU. The moment one request exists, the loop spins without ever sleeping until every request is finished. An idle vLLM server shows 0% CPU for the engine core; a loaded one shows 100% of one core. Both are correct.

One step, the simple way#

EngineCore.step() is the version to learn first:

Python
def step(self) -> tuple[dict[int, EngineCoreOutputs], bool]:
    if not self.scheduler.has_requests():
        return {}, False
    scheduler_output = self.scheduler.schedule(self._should_throttle_prefills())
    future = self.model_executor.execute_model(scheduler_output, non_block=True)
    grammar_output = self.scheduler.get_grammar_bitmask(scheduler_output)
    with (...):
        model_output = future.result()
        if model_output is None:
            model_output = self.model_executor.sample_tokens(grammar_output)

    # Before processing the model output, process any aborts that happened
    # during the model execution.
    self._process_aborts_queue()
    engine_core_outputs = self.scheduler.update_from_output(
        scheduler_output, model_output
    )
    return engine_core_outputs, scheduler_output.total_num_scheduled_tokens > 0

Five things happen, in this order:

#CallWhat it doesRuns on
1scheduler.schedule()Chooses how many tokens of which requests run; allocates KV blocks; returns a SchedulerOutputCPU, engine core
2execute_model(non_block=True)Ships the SchedulerOutput to the workers and returns at once with a future—
3get_grammar_bitmask()For requests using structured output, computes which tokens the grammar allows nextCPU, while the GPU runs
4future.result() then sample_tokens()Waits for the forward pass, then sends the bitmask and asks the workers to sampleGPU
5update_from_output()Appends each new token to its request, checks stop conditions, frees blocks of finished requests, builds EngineCoreOutputsCPU, engine core

Why the model call is split in two. The forward pass does not need to know which tokens a JSON grammar permits; only sampling does. So execute_model starts the forward pass and returns None, the engine core computes the grammar bitmask on the CPU during the forward pass, and sample_tokens(grammar_output) delivers it just in time. Structured output therefore costs CPU time that mostly overlaps with GPU time rather than adding to it.

Why aborts are processed between 4 and 5. A forward pass takes milliseconds to hundreds of milliseconds. Any client that disconnected during it has an abort waiting. Applying those aborts before update_from_output means the scheduler never appends a token to — or allocates another block for — a request nobody is waiting for.

One step, the default way: the batch queue#

step() has a flaw. Between “forward pass finished” and “next forward pass started”, the GPU is idle while Python runs update_from_output and then schedule. That is typically one to a few milliseconds, on every token.

So the engine keeps a small queue of batches that have been scheduled but not finished:

Python
@property
def max_concurrent_batches(self) -> int:
    # PP requires PP-size concurrent batches to fill the pipeline.
    # Async scheduling requires 2 concurrent batches to overlap.
    pp_size = self.parallel_config.pipeline_parallel_size
    if self.scheduler_config.async_scheduling:
        if self.use_v2_model_runner:
            return pp_size + 1
        if pp_size <= 1:
            return 2
    return pp_size

Async scheduling is on by default (it is disabled automatically for embedding models and for a few speculative-decoding methods). The queue size is therefore 2 on an ordinary single-GPU server, and the engine uses step_with_batch_queue instead of step:

1. If the queue is not full and the scheduler has requests:
       schedule a new batch, start it (non-blocking), push it on the queue.
       If the queue still has room, return immediately with no output.
2. Otherwise block on the OLDEST batch in the queue.
3. update_from_output() for that batch; return its outputs.

The source comment states the priority: “fulfilling the batch queue has a higher priority than getting model outputs.”

The effect is that batch N+1 is scheduled and sent to the worker while batch N is still on the GPU. When N finishes, N+1 is already waiting in the worker and starts with no gap.

synchronous        GPU: [ step 1 ]......[ step 2 ]......[ step 3 ]
                   CPU:           [u+s]          [u+s]

async scheduling   GPU: [ step 1 ][ step 2 ][ step 3 ][ step 4 ]
                   CPU: [s2]  [u1+s3] [u2+s4] [u3+s5]
                                                     u = update_from_output, s = schedule

There is a catch, and it shapes the scheduler. When batch N+1 is being scheduled, the tokens that batch N will sample do not exist yet. The scheduler has to schedule the next decode step for a request without knowing its previous output token. It does this with placeholders; the worker, which does know the token, fills it in. That mechanism is the subject of Async Scheduling.

Pipeline parallelism uses the same queue for a different reason. With the model split across two GPUs in series, GPU 0 is free as soon as it has handed its activations to GPU 1. A queue of size PP keeps every stage busy with a different batch, removing what the code calls “pipeline bubbles”.

What happens after a step#

Python
def _process_engine_step(self) -> bool:
    outputs, model_executed = self.step_fn()
    for output in outputs.items() if outputs else ():
        self.output_queue.put_nowait(output)
    self.post_step(model_executed)
    # If no model execution happened but there is still scheduler work
    # (e.g. WAITING_FOR_REMOTE_KVS or delayed KV connector frees), yield
    # the GIL briefly to allow background transfer threads to make progress.
    if not model_executed and self.scheduler.has_requests():
        time.sleep(0.001)
    return model_executed
  • Outputs are keyed by client index: a step’s results are split per API server, so each API server receives only its own requests.
  • post_step collects draft tokens for speculative decoding when async scheduling is off.
  • The one-millisecond sleep is the single place the busy loop yields. It applies when requests exist but none could run — for instance when every waiting request is waiting for KV data to arrive from another machine. Without it the loop would spin at 100% CPU and starve the threads doing that transfer.

Messages other than requests#

_handle_client_request dispatches on the type byte from Processes and Wires:

TypeAction
ADDscheduler.add_request(request) — the request joins the waiting queue
ABORTscheduler.finish_requests(ids, FINISHED_ABORTED) — idempotent, so the duplicate on the fast aborts queue is harmless
UTILITYLook up a method on the engine core by name, call it, post the result
EXECUTOR_FAILEDraise RuntimeError("Executor failed.") — a worker died; take the engine down
WAKEUPNothing. It exists only to unblock input_queue.get()

Utility calls run on the main loop thread, between steps. That is why operations such as loading a LoRA adapter or resetting the prefix cache are safe without locks: no step is in progress while they run. It is also why a slow utility call stalls generation for everyone.

Pause, sleep and wake#

Three utility methods change the loop’s behaviour without stopping the process.

pause_scheduler(mode) has three modes:

ModeIn-flight requestsNew requests
abort (default)Aborted immediatelyQueued until resume
waitAllowed to finishQueued until resume
keepFrozen in place; no steps runQueued until resume

sleep(level) builds on it, to release GPU memory while keeping the process alive:

LevelEffectUse
0Pause scheduling only. No memory changes.A brief hold
1Move weights to CPU RAM, discard the KV cacheShare a GPU between models that take turns; waking is fast
2Discard all GPU memoryLongest idle; waking reloads weights from disk

wake_up() reverses it and resumes scheduling once all memory is back. The server exposes these as /sleep, /wake_up and /is_sleeping when started with --enable-sleep-mode. Reinforcement-learning loops use pause with wait to swap in new weights between batches of generation.

Start-up of the engine core#

Before the loop runs, EngineCore.__init__ does the work whose results you read in the startup log:

  1. Create the executor, which starts the workers and loads the model.
  2. _initialize_kv_caches: ask the workers for each layer’s cache requirements, profile free memory, compute the number of blocks, tell the workers to allocate them, then compile and warm up (Sizing the Cache).
  3. Create the structured output manager and the scheduler.
  4. Choose step or step_with_batch_queue.
  5. freeze_gc_heap(): mark everything allocated so far as permanent so the garbage collector never scans it again.

Two decisions are taken here that surprise people later: if the model has no KV cache (an encoder-only embedding model), chunked prefill is switched off; if any layer attends non-causally, both chunked prefill and prefix caching are switched off, because both assume a token’s KV depends only on the tokens before it.

Code#

How much does overlapping scheduling with the GPU buy? This program times both loops for a range of batch sizes. CPU cost per step grows with the number of requests; GPU cost grows more slowly, so the idle gap matters most exactly when the server is busy.

Go
package main

import "fmt"

func main() {
	fmt.Printf("%8s %10s %10s %14s %14s %9s\n",
		"requests", "cpu ms", "gpu ms", "sync tok/s", "async tok/s", "gain")

	for _, n := range []int{1, 8, 32, 128, 256, 512} {
		// Rough shapes, not measurements: Python work is linear in batch size,
		// a decode forward pass has a fixed cost plus a small per-request cost.
		cpu := 0.25 + 0.012*float64(n) // schedule + update_from_output
		gpu := 8.0 + 0.03*float64(n)   // one decode step

		syncStep := cpu + gpu // the GPU waits while Python plans the next step
		asyncStep := gpu      // planning overlaps the previous forward pass
		if cpu > gpu {
			asyncStep = cpu // Python is now the bottleneck
		}

		syncTok := float64(n) / syncStep * 1000
		asyncTok := float64(n) / asyncStep * 1000
		fmt.Printf("%8d %10.2f %10.2f %14.0f %14.0f %8.1f%%\n",
			n, cpu, gpu, syncTok, asyncTok, 100*(asyncTok/syncTok-1))
	}
}

The gain grows with concurrency, which is why the feature became the default. Notice the if cpu > gpu branch as well: once Python’s per-step work exceeds the forward pass, more GPU does not help. That happens with very small models at very high concurrency, and it is the motivation for the Rust frontend and for moving work out of the engine core’s main thread.

Remember this#

  • The engine core is a loop: drain inbox, schedule, execute, update, post.
  • It blocks only when it has no requests; otherwise it uses one full core.
  • The model call is split so grammar bitmasks are computed while the GPU runs.
  • Aborts are applied right after each forward pass, before outputs are processed.
  • By default a batch queue of size 2 lets step N+1 be scheduled while step N is on the GPU.
  • Utility calls (LoRA, cache reset, sleep) run between steps on the loop thread.
  • sleep levels 0, 1, 2 trade wake-up time for released GPU memory.

Try it#

  1. In the program, change the GPU cost to 2.0 + 0.01*n (a much smaller model). At which batch size does Python become the bottleneck?
  2. On a running server, watch the engine core’s CPU with top while idle and while serving one long request. Explain both readings using has_work().
  3. Start a server with --enable-sleep-mode, POST to /sleep?level=1, and watch nvidia-smi. Then POST to /wake_up. How long does each take, and where did the weights go in between?

Check yourself#

  1. In step(), why is get_grammar_bitmask called after execute_model and not before it?
  2. What information is missing when the engine schedules batch N+1 before batch N has finished?
  3. Why is it safe to add a LoRA adapter without a lock?

Sources#

Checked on 5 October 2026 against main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom