The idea in one minute#
The OpenAI chat API returns three separate things: content, reasoning and tool_calls. A
language model returns one thing: a stream of tokens. The model was trained to mark the parts
itself — its thinking between special tags, its tool calls in some agreed JSON-like syntax —
and every model family chose different markers. Tool calling and reasoning in vLLM are
therefore almost entirely a text-processing feature that lives in the API server: the chat
template writes the tool definitions into the prompt in the model’s format, and a per-family
parser splits the output stream back into the three fields. The engine core is not
involved, except in the cases where the output must be forced into shape.
A picture#
flowchart LR REQ[":i-file-text: <b>Request</b><br/><small>messages + tools[]</small>"] --> TPL[":i-code: <b>Chat template</b><br/><small>renders tools into the prompt<br/>in this model's syntax</small>"] TPL --> ENG[":vllm: <b>Engine</b><br/><small>tokens in, tokens out</small>"] ENG --> RP[":i-brain: <b>Reasoning parser</b><br/><small>splits thinking from answer</small>"] RP -->|"reasoning"| OUT RP --> TP[":i-wrench: <b>Tool parser</b><br/><small>finds and extracts calls</small>"] TP -->|"content"| OUT[":i-package: <b>Response</b><br/><small>reasoning, content, tool_calls</small>"] TP -->|"tool_calls"| OUT SO[":i-shield-check: <b>Structured output</b><br/><small>only for required / named choice</small>"] -.->|"grammar"| ENG class REQ,OUT neutral class TPL,RP,TP compute class ENG queue class SO warn
How it really works#
The way in: tools are prompt text#
A request’s tools array is a list of JSON schemas. There is no separate channel for it into
the model. The renderer passes it to the chat template, which writes it into the prompt in
whatever layout this model was trained on: a JSON list inside special tags for one family, a
block of function signatures for another.
Two consequences:
- Tool definitions cost tokens, on every request. Twenty tools with detailed descriptions can be several thousand tokens. They are identical across requests, so place them where the prefix cache can absorb them, and keep their order stable.
- The right template is essential. A model whose default template ignores
toolswill never call one. Some models need a template shipped in vLLM’sexamples/directory, passed with--chat-template.
The way out: tool_choice decides the mechanism#
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--enable-auto-tool-choice \
--tool-call-parser llama3_json \
--chat-template examples/tool_chat_template_llama3.1_json.jinjatool_choice | Meaning | How vLLM produces the call |
|---|---|---|
"none" | Never call a tool | Tools may still be shown to the model, depending on configuration; output is plain content |
"auto" | The model decides | The model writes a call in its own syntax; a parser recognises and extracts it. Needs --enable-auto-tool-choice and --tool-call-parser. |
"required" | Must call at least one tool | Structured output forces a JSON array of calls matching the tools’ schemas |
{"type": "function", "function": {"name": "f"}} | Must call f | Structured output forces JSON matching f’s parameter schema |
The last two rows use the grammar machinery from Sampling and Structured Output: the arguments are guaranteed to be valid JSON matching the schema, because invalid tokens are masked. No parser is needed.
auto is different in kind. The model is free; it may answer in prose, call one tool, or call
several. Nothing constrains its output, so the arguments are only as well-formed as the model
makes them. The parser’s job is to find calls in free text.
Tool parsers#
A tool parser (vllm/tool_parsers/) is a class for one output syntax. There are dozens:
hermes, llama3_json, mistral, pythonic, deepseek_v3 and more. Each implements two
methods:
| Method | Used for | Difficulty |
|---|---|---|
| Extract from complete text | Non-streaming responses | Easy: the whole output is available |
| Extract from a stream | Streaming responses | Hard: decide, token by token, whether text is content, the start of a marker, or part of a call |
The streaming problem is the stop-string problem from
AsyncLLM and the Way Back with more
state. If the output so far ends in <tool, the parser cannot emit it as content — it may be
the start of <tool_call> — and cannot treat it as a call yet. It holds the text back until
the next token settles the question. Streamed tool-call arguments are emitted as
incremental JSON fragments, as the OpenAI protocol requires.
Choosing the wrong parser is the most common tool-calling failure. The symptom is a response
whose content contains raw <tool_call>{...} text and whose tool_calls is empty: the model
did its part and nobody translated it.
parallel_tool_calls: false asks vLLM to return at most one call per response.
Reasoning parsers#
Reasoning models emit their thinking before the answer, delimited by markers such as
<think>…</think>. --reasoning-parser <name> selects a parser that moves that text into a
separate field:
{"message": {"role": "assistant",
"reasoning": "The user wants the weather, so…",
"content": "It is 31 °C in Pune."}}warning
The field is called reasoning. Older versions of vLLM called it reasoning_content. The
documentation warns that a client still reading the old name “could silently read an empty
reasoning_content, even when reasoning is populated.”
Three details that cause surprises:
- Reasoning counts as output. It consumes
max_tokensand is billed, timed and scheduled like any other generated token. A request withmax_tokens: 200to a model that thinks for 300 tokens returns no answer at all. - The opening marker may be in the prompt. Many templates end the prompt with the opening tag, so the model’s output starts mid-thought with no marker to find. Each parser knows whether its family does this; that is one reason parsers are per-family.
- Past reasoning is usually dropped from history. Templates commonly omit earlier turns' reasoning when rendering the next prompt, which changes the token sequence and ends the prefix-cache hit at that point. Some models and templates support keeping it (“interleaved thinking”) for multi-step tool use.
Where the two meet structured output#
A grammar says “the output is JSON matching this schema”. A reasoning model’s output is thinking, then JSON. Applying the grammar from the first token would forbid the thinking.
So the structured-output manager is told which reasoning parser is active
(StructuredOutputsConfig.reasoning_parser) and uses it to find where reasoning ends. The
grammar is switched on only from that point. The request carries a reasoning_ended flag for
this, and the bitmask is simply not applied while it is false.
The documentation’s table of reasoning models has a column for which structured-output modes and whether tool calling work with each parser. Check it before combining the three on a new model; not every combination is supported.
Why this lives in the API server#
Parsers need text, and only the API server has a tokenizer and the detokenised stream. Running them there keeps the engine core free of per-model string handling and keeps its main thread for scheduling.
It also means the parsers’ CPU cost scales with the API server, not the engine. For
tool-heavy, high-throughput traffic the fix for parser overhead is more API server processes,
or the Rust frontend, which reimplements template rendering and reasoning and tool parsing in
its vllm-chat crate
(The Rust Frontend and What Is Next).
Adding your own#
Both kinds of parser are pluggable. --tool-parser-plugin path/to/file.py loads a module that
registers a new tool parser; --reasoning-parser-plugin does the same for reasoning. This is
how support for a newly released model family is added before it lands upstream
(Plugins).
A checklist for a model that will not call tools#
- Does the rendered prompt contain the tool definitions? Call
/tokenizewith yourmessagesandtools, then/detokenize, and read it. - Is
--enable-auto-tool-choiceset, with the--tool-call-parserdocumented for this model? - Does the raw output contain a call in a syntax the parser does not expect? Send the request
without the parser and read
content. - Is
max_tokenslarge enough for reasoning plus the call? - If arguments are malformed under
auto, tryrequiredor a named tool: the grammar then guarantees the JSON.
Code#
A streaming parser that separates reasoning, content and a tool call, holding back text that might be the beginning of a marker.
package main
import (
"fmt"
"strings"
)
const (
thinkOpen, thinkClose = "<think>", "</think>"
toolOpen, toolClose = "<tool_call>", "</tool_call>"
)
// parser splits a token stream into three channels, as vLLM's reasoning and
// tool parsers do for a streaming chat completion.
type parser struct {
buf string // text not yet classified
state string // "content", "reasoning" or "tool"
}
type delta struct{ kind, text string }
// heldBack returns how many trailing bytes of s could be the start of marker.
func heldBack(s, marker string) int {
for n := min(len(marker)-1, len(s)); n > 0; n-- {
if strings.HasPrefix(marker, s[len(s)-n:]) {
return n
}
}
return 0
}
func (p *parser) feed(piece string) []delta {
p.buf += piece
var out []delta
for {
var open, close, inner string
switch p.state {
case "content":
// Either marker may come next; take whichever appears first.
ti, ki := strings.Index(p.buf, thinkOpen), strings.Index(p.buf, toolOpen)
switch {
case ti >= 0 && (ki < 0 || ti < ki):
open, inner = thinkOpen, "reasoning"
case ki >= 0:
open, inner = toolOpen, "tool"
}
if open != "" {
i := strings.Index(p.buf, open)
if i > 0 {
out = append(out, delta{"content", p.buf[:i]})
}
p.buf = p.buf[i+len(open):]
p.state = inner
continue
}
// No marker yet: release everything that cannot be the start of one.
hold := max(heldBack(p.buf, thinkOpen), heldBack(p.buf, toolOpen))
if emit := p.buf[:len(p.buf)-hold]; emit != "" {
out = append(out, delta{"content", emit})
}
p.buf = p.buf[len(p.buf)-hold:]
return out
case "reasoning":
close = thinkClose
case "tool":
close = toolClose
}
if i := strings.Index(p.buf, close); i >= 0 {
if i > 0 {
out = append(out, delta{p.state, p.buf[:i]})
}
p.buf = p.buf[i+len(close):]
p.state = "content"
continue
}
if p.state == "reasoning" { // reasoning streams as it arrives
hold := heldBack(p.buf, close)
if emit := p.buf[:len(p.buf)-hold]; emit != "" {
out = append(out, delta{"reasoning", emit})
}
p.buf = p.buf[len(p.buf)-hold:]
}
// A tool call is held until its closing tag: partial JSON is not yet a call.
return out
}
}
func main() {
stream := []string{"<th", "ink>", "The user wants", " weather.", "</thi", "nk>",
"Let me", " check.", "<tool", "_call>", `{"name": "get_weather",`,
` "arguments": {"city": "Pune"}}`, "</tool_call>"}
p := &parser{state: "content"}
channels := map[string]string{}
for i, piece := range stream {
ds := p.feed(piece)
fmt.Printf("token %-2d %-34q", i+1, piece)
if len(ds) == 0 {
fmt.Print(" (held back)")
}
for _, d := range ds {
fmt.Printf(" %s+=%q", d.kind, d.text)
channels[d.kind] += d.text
}
fmt.Println()
}
fmt.Println()
fmt.Printf("message.reasoning : %q\n", channels["reasoning"])
fmt.Printf("message.content : %q\n", channels["content"])
fmt.Printf("message.tool_calls: %s\n", channels["tool"])
fmt.Println(`finish_reason : "tool_calls"`)
}Eight of the thirteen tokens produce no output when they arrive. Some are fragments of a marker and cannot be classified yet; the rest belong to a tool call that is withheld until its closing tag, because half a JSON object is not a call. A real parser streams argument fragments rather than waiting, which requires tracking JSON nesting as well; the hold-back principle is the same.
Remember this#
- Tools go in as prompt text through the chat template, and cost tokens on every request.
autolets the model write calls in its own syntax and relies on a per-family parser.requiredand named tool choice use structured output; the arguments are guaranteed valid JSON.- Reasoning is split into a separate
reasoningfield by a per-family parser, and counts againstmax_tokens. - With a reasoning model, the grammar is enabled only after reasoning ends.
- Parsers run in the API server; their cost scales with API server processes.
- Raw call syntax appearing in
contentmeans the wrong parser, or none, is configured.
Try it#
- Add a second tool call after the first in
stream. Does the parser return both? What shouldfinish_reasonbe? - Remove the
<think>opening tag from the stream, as if the template had put it in the prompt. What does the parser do, and how would you fix it for such a model? - On a real server, send the same tool-calling request with
tool_choiceset to"auto"and to"required". Compare latency and inspect the arguments’ JSON in each case.
Check yourself#
- How do tool definitions reach the model?
- What is the mechanical difference between
tool_choice: "auto"andtool_choice: "required"? - Why must a streaming parser sometimes hold back text it has already received?
Sources#
Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.
- Tool calling
- Reasoning outputs — including the
reasoningfield rename and the support table - Interleaved thinking
vllm/tool_parsers/andvllm/reasoning/vllm/config/structured_outputs.py—reasoning_parser