Pidoku

Tool Calling and Reasoning

Advanced 40 min Difficulty 3/5 Lesson 04 of 05

Prerequisites The Frontend, AsyncLLM and the Way Back, Sampling and Structured Output

The idea in one minute#

The OpenAI chat API returns three separate things: content, reasoning and tool_calls. A language model returns one thing: a stream of tokens. The model was trained to mark the parts itself — its thinking between special tags, its tool calls in some agreed JSON-like syntax — and every model family chose different markers. Tool calling and reasoning in vLLM are therefore almost entirely a text-processing feature that lives in the API server: the chat template writes the tool definitions into the prompt in the model’s format, and a per-family parser splits the output stream back into the three fields. The engine core is not involved, except in the cases where the output must be forced into shape.

A picture#

flowchart LR
  REQ[":i-file-text: <b>Request</b><br/><small>messages + tools[]</small>"] --> TPL[":i-code: <b>Chat template</b><br/><small>renders tools into the prompt<br/>in this model's syntax</small>"]
  TPL --> ENG[":vllm: <b>Engine</b><br/><small>tokens in, tokens out</small>"]
  ENG --> RP[":i-brain: <b>Reasoning parser</b><br/><small>splits thinking from answer</small>"]
  RP -->|"reasoning"| OUT
  RP --> TP[":i-wrench: <b>Tool parser</b><br/><small>finds and extracts calls</small>"]
  TP -->|"content"| OUT[":i-package: <b>Response</b><br/><small>reasoning, content, tool_calls</small>"]
  TP -->|"tool_calls"| OUT
  SO[":i-shield-check: <b>Structured output</b><br/><small>only for required / named choice</small>"] -.->|"grammar"| ENG
  class REQ,OUT neutral
  class TPL,RP,TP compute
  class ENG queue
  class SO warn

How it really works#

The way in: tools are prompt text#

A request’s tools array is a list of JSON schemas. There is no separate channel for it into the model. The renderer passes it to the chat template, which writes it into the prompt in whatever layout this model was trained on: a JSON list inside special tags for one family, a block of function signatures for another.

Two consequences:

  • Tool definitions cost tokens, on every request. Twenty tools with detailed descriptions can be several thousand tokens. They are identical across requests, so place them where the prefix cache can absorb them, and keep their order stable.
  • The right template is essential. A model whose default template ignores tools will never call one. Some models need a template shipped in vLLM’s examples/ directory, passed with --chat-template.

The way out: tool_choice decides the mechanism#

Shell
vllm serve meta-llama/Llama-3.1-8B-Instruct \
    --enable-auto-tool-choice \
    --tool-call-parser llama3_json \
    --chat-template examples/tool_chat_template_llama3.1_json.jinja
tool_choiceMeaningHow vLLM produces the call
"none"Never call a toolTools may still be shown to the model, depending on configuration; output is plain content
"auto"The model decidesThe model writes a call in its own syntax; a parser recognises and extracts it. Needs --enable-auto-tool-choice and --tool-call-parser.
"required"Must call at least one toolStructured output forces a JSON array of calls matching the tools’ schemas
{"type": "function", "function": {"name": "f"}}Must call fStructured output forces JSON matching f’s parameter schema

The last two rows use the grammar machinery from Sampling and Structured Output: the arguments are guaranteed to be valid JSON matching the schema, because invalid tokens are masked. No parser is needed.

auto is different in kind. The model is free; it may answer in prose, call one tool, or call several. Nothing constrains its output, so the arguments are only as well-formed as the model makes them. The parser’s job is to find calls in free text.

Tool parsers#

A tool parser (vllm/tool_parsers/) is a class for one output syntax. There are dozens: hermes, llama3_json, mistral, pythonic, deepseek_v3 and more. Each implements two methods:

MethodUsed forDifficulty
Extract from complete textNon-streaming responsesEasy: the whole output is available
Extract from a streamStreaming responsesHard: decide, token by token, whether text is content, the start of a marker, or part of a call

The streaming problem is the stop-string problem from AsyncLLM and the Way Back with more state. If the output so far ends in <tool, the parser cannot emit it as content — it may be the start of <tool_call> — and cannot treat it as a call yet. It holds the text back until the next token settles the question. Streamed tool-call arguments are emitted as incremental JSON fragments, as the OpenAI protocol requires.

Choosing the wrong parser is the most common tool-calling failure. The symptom is a response whose content contains raw <tool_call>{...} text and whose tool_calls is empty: the model did its part and nobody translated it.

parallel_tool_calls: false asks vLLM to return at most one call per response.

Reasoning parsers#

Reasoning models emit their thinking before the answer, delimited by markers such as <think>…</think>. --reasoning-parser <name> selects a parser that moves that text into a separate field:

JSON
{"message": {"role": "assistant",
             "reasoning": "The user wants the weather, so…",
             "content": "It is 31 °C in Pune."}}

warning

The field is called reasoning. Older versions of vLLM called it reasoning_content. The documentation warns that a client still reading the old name “could silently read an empty reasoning_content, even when reasoning is populated.”

Three details that cause surprises:

  • Reasoning counts as output. It consumes max_tokens and is billed, timed and scheduled like any other generated token. A request with max_tokens: 200 to a model that thinks for 300 tokens returns no answer at all.
  • The opening marker may be in the prompt. Many templates end the prompt with the opening tag, so the model’s output starts mid-thought with no marker to find. Each parser knows whether its family does this; that is one reason parsers are per-family.
  • Past reasoning is usually dropped from history. Templates commonly omit earlier turns' reasoning when rendering the next prompt, which changes the token sequence and ends the prefix-cache hit at that point. Some models and templates support keeping it (“interleaved thinking”) for multi-step tool use.

Where the two meet structured output#

A grammar says “the output is JSON matching this schema”. A reasoning model’s output is thinking, then JSON. Applying the grammar from the first token would forbid the thinking.

So the structured-output manager is told which reasoning parser is active (StructuredOutputsConfig.reasoning_parser) and uses it to find where reasoning ends. The grammar is switched on only from that point. The request carries a reasoning_ended flag for this, and the bitmask is simply not applied while it is false.

The documentation’s table of reasoning models has a column for which structured-output modes and whether tool calling work with each parser. Check it before combining the three on a new model; not every combination is supported.

Why this lives in the API server#

Parsers need text, and only the API server has a tokenizer and the detokenised stream. Running them there keeps the engine core free of per-model string handling and keeps its main thread for scheduling.

It also means the parsers’ CPU cost scales with the API server, not the engine. For tool-heavy, high-throughput traffic the fix for parser overhead is more API server processes, or the Rust frontend, which reimplements template rendering and reasoning and tool parsing in its vllm-chat crate (The Rust Frontend and What Is Next).

Adding your own#

Both kinds of parser are pluggable. --tool-parser-plugin path/to/file.py loads a module that registers a new tool parser; --reasoning-parser-plugin does the same for reasoning. This is how support for a newly released model family is added before it lands upstream (Plugins).

A checklist for a model that will not call tools#

  1. Does the rendered prompt contain the tool definitions? Call /tokenize with your messages and tools, then /detokenize, and read it.
  2. Is --enable-auto-tool-choice set, with the --tool-call-parser documented for this model?
  3. Does the raw output contain a call in a syntax the parser does not expect? Send the request without the parser and read content.
  4. Is max_tokens large enough for reasoning plus the call?
  5. If arguments are malformed under auto, try required or a named tool: the grammar then guarantees the JSON.

Code#

A streaming parser that separates reasoning, content and a tool call, holding back text that might be the beginning of a marker.

Go
package main

import (
	"fmt"
	"strings"
)

const (
	thinkOpen, thinkClose = "<think>", "</think>"
	toolOpen, toolClose   = "<tool_call>", "</tool_call>"
)

// parser splits a token stream into three channels, as vLLM's reasoning and
// tool parsers do for a streaming chat completion.
type parser struct {
	buf   string // text not yet classified
	state string // "content", "reasoning" or "tool"
}

type delta struct{ kind, text string }

// heldBack returns how many trailing bytes of s could be the start of marker.
func heldBack(s, marker string) int {
	for n := min(len(marker)-1, len(s)); n > 0; n-- {
		if strings.HasPrefix(marker, s[len(s)-n:]) {
			return n
		}
	}
	return 0
}

func (p *parser) feed(piece string) []delta {
	p.buf += piece
	var out []delta
	for {
		var open, close, inner string
		switch p.state {
		case "content":
			// Either marker may come next; take whichever appears first.
			ti, ki := strings.Index(p.buf, thinkOpen), strings.Index(p.buf, toolOpen)
			switch {
			case ti >= 0 && (ki < 0 || ti < ki):
				open, inner = thinkOpen, "reasoning"
			case ki >= 0:
				open, inner = toolOpen, "tool"
			}
			if open != "" {
				i := strings.Index(p.buf, open)
				if i > 0 {
					out = append(out, delta{"content", p.buf[:i]})
				}
				p.buf = p.buf[i+len(open):]
				p.state = inner
				continue
			}
			// No marker yet: release everything that cannot be the start of one.
			hold := max(heldBack(p.buf, thinkOpen), heldBack(p.buf, toolOpen))
			if emit := p.buf[:len(p.buf)-hold]; emit != "" {
				out = append(out, delta{"content", emit})
			}
			p.buf = p.buf[len(p.buf)-hold:]
			return out
		case "reasoning":
			close = thinkClose
		case "tool":
			close = toolClose
		}

		if i := strings.Index(p.buf, close); i >= 0 {
			if i > 0 {
				out = append(out, delta{p.state, p.buf[:i]})
			}
			p.buf = p.buf[i+len(close):]
			p.state = "content"
			continue
		}
		if p.state == "reasoning" { // reasoning streams as it arrives
			hold := heldBack(p.buf, close)
			if emit := p.buf[:len(p.buf)-hold]; emit != "" {
				out = append(out, delta{"reasoning", emit})
			}
			p.buf = p.buf[len(p.buf)-hold:]
		}
		// A tool call is held until its closing tag: partial JSON is not yet a call.
		return out
	}
}

func main() {
	stream := []string{"<th", "ink>", "The user wants", " weather.", "</thi", "nk>",
		"Let me", " check.", "<tool", "_call>", `{"name": "get_weather",`,
		` "arguments": {"city": "Pune"}}`, "</tool_call>"}

	p := &parser{state: "content"}
	channels := map[string]string{}
	for i, piece := range stream {
		ds := p.feed(piece)
		fmt.Printf("token %-2d %-34q", i+1, piece)
		if len(ds) == 0 {
			fmt.Print(" (held back)")
		}
		for _, d := range ds {
			fmt.Printf(" %s+=%q", d.kind, d.text)
			channels[d.kind] += d.text
		}
		fmt.Println()
	}
	fmt.Println()
	fmt.Printf("message.reasoning : %q\n", channels["reasoning"])
	fmt.Printf("message.content   : %q\n", channels["content"])
	fmt.Printf("message.tool_calls: %s\n", channels["tool"])
	fmt.Println(`finish_reason     : "tool_calls"`)
}

Eight of the thirteen tokens produce no output when they arrive. Some are fragments of a marker and cannot be classified yet; the rest belong to a tool call that is withheld until its closing tag, because half a JSON object is not a call. A real parser streams argument fragments rather than waiting, which requires tracking JSON nesting as well; the hold-back principle is the same.

Remember this#

  • Tools go in as prompt text through the chat template, and cost tokens on every request.
  • auto lets the model write calls in its own syntax and relies on a per-family parser.
  • required and named tool choice use structured output; the arguments are guaranteed valid JSON.
  • Reasoning is split into a separate reasoning field by a per-family parser, and counts against max_tokens.
  • With a reasoning model, the grammar is enabled only after reasoning ends.
  • Parsers run in the API server; their cost scales with API server processes.
  • Raw call syntax appearing in content means the wrong parser, or none, is configured.

Try it#

  1. Add a second tool call after the first in stream. Does the parser return both? What should finish_reason be?
  2. Remove the <think> opening tag from the stream, as if the template had put it in the prompt. What does the parser do, and how would you fix it for such a model?
  3. On a real server, send the same tool-calling request with tool_choice set to "auto" and to "required". Compare latency and inspect the arguments’ JSON in each case.

Check yourself#

  1. How do tool definitions reach the model?
  2. What is the mechanical difference between tool_choice: "auto" and tool_choice: "required"?
  3. Why must a streaming parser sometimes hold back text it has already received?

Sources#

Checked on 5 October 2026 against vLLM v0.30.0 and main at commit 0c16eee.

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom