The idea in one minute#
Applications should never call a model directly. They call an AI gateway, and the gateway calls models. That one indirection is where every cross-cutting concern lives: who is calling, how many tokens they may spend, which model serves this request, what happens when that model is slow, and what gets logged. Without it, each of those is reimplemented — differently — in every service.
An AI gateway differs from an ordinary API gateway in three ways: it meters tokens, it routes on model state, and it handles streams that last half a minute.
A picture#
flowchart LR
APP[":i-code: <b>Applications</b><br/><small>one API, one key each</small>"] --> GW
subgraph GW["AI gateway"]
direction TB
A1[":i-fingerprint: Authenticate, attribute"] --> A2[":i-coins: Token limits and budgets"]
A2 --> A3[":i-route: Route by model, cost, health"]
A3 --> A4[":i-shield-check: Guardrail hooks"]
end
GW --> P1[":anthropic: <b>Hosted API A</b>"]
GW --> P2[":googlegemini: <b>Hosted API B</b>"]
GW --> P3[":vllm: <b>Your own pool</b><br/><small>InferencePool on Kubernetes</small>"]
P1 -.->|"slow or failing"| P2
GW -.-> T[":opentelemetry: <b>Usage events</b><br/><small>tokens, latency, tenant</small>"]
class APP neutral
class A1,A2,A3 queue
class A4 warn
class P1,P2 io
class P3 compute
class T memoryHow it really works#
Hosted, managed or self-hosted#
The first decision is where the model runs. It is a ladder; climb only when a reason forces you.
| Rung | You operate | Choose it when | You give up |
|---|---|---|---|
| Hosted API | Nothing | Starting out; you need frontier quality; volume is modest or spiky | Control of latency, data location, price |
| Dedicated capacity from a provider | Nothing, but you reserve throughput | Latency must be predictable; volume is steady | Flexibility; you pay for idle |
| Managed open-weights endpoint | Configuration | You need a specific open model or fine-tune without running GPUs | Some efficiency and tuning freedom |
| Self-hosted on your GPUs | Everything | Data cannot leave; volume is large and steady; you need custom models or the lowest unit cost | Simplicity. This is a platform team |
Most real systems use two rungs at once: a frontier API for the hard calls and a self-hosted or small model for the many easy ones. The gateway is what makes that mixture invisible to application code. AI Infrastructure From Scratch covers the bottom rung.
What the gateway does#
Identity and attribution. Applications hold a gateway key, never a provider key. Each request is attributed to a tenant, application and user, and that attribution rides on every usage event. Provider credentials exist in exactly one place.
Token-based limits. Limits are in tokens per minute and dollars per day, per tenant.
There is a wrinkle: output tokens are not known when the request arrives. Gateways handle this
by reserving an estimate (input tokens plus max_tokens) and settling the difference when the
response finishes.
Routing. A request names a logical model — “fast”, “smart”, “code” — and the gateway maps it to a real one. Routing can consider price, health, the tenant’s data-residency rule and, for your own pool, live engine state: the Gateway API Inference Extension routes to the replica with the shortest queue and the warmest cache rather than round-robin.
Fallback. When a route errors or exceeds its deadline for first token, the gateway retries on another. Two rules keep fallback safe: fall back only before the first token has been sent to the client, and send fallback traffic to a route with spare capacity — otherwise one provider’s outage becomes two.
Caching. Provider-side prompt caching needs stable prefixes, and the gateway is where you can enforce prompt layout and measure the cache hit rate. Exact-match response caching is safe for deterministic calls; semantic caching is covered, with its risks, in State, Memory and Caching.
Guardrail hooks. The gateway is the natural place to call input and output checks, because every request already passes through it. What those checks can and cannot do is the subject of Runtime Guardrails.
One record per request. Tenant, model, route, input, cached and output tokens, TTFT, total time, finish reason, cost. This single wide event is the source of truth for billing, capacity planning and incident review.
Model routing as a cost lever#
Not every call needs the largest model. A router that sends easy calls to a small model cuts cost several-fold, and the design question is how it decides:
| Strategy | How | Trade-off |
|---|---|---|
| By task | The application says which tier each call site needs | Simple and predictable; needs discipline |
| Cascade | Try the small model; escalate when a check fails | Saves most; adds latency on escalation |
| Learned router | A classifier predicts which model is sufficient | Best savings; another model to evaluate |
Whatever the strategy, changes to routing are behaviour changes and go through the same evaluation gate as a model upgrade.
What fills the box in 2026#
| Tool | Notes |
|---|---|
| Envoy AI Gateway | Built on Envoy Gateway; 1.0 in June 2026. Token-based rate limiting, provider failover, MCP routing, OpenTelemetry GenAI tracing |
| agentgateway | Rust data plane aimed at agent traffic: LLM, MCP and A2A in one proxy; integrates with InferencePool routing |
| LiteLLM | Python proxy with the widest provider coverage; quick to adopt |
| Kong, cloud provider gateways | Sensible when you already run them for other APIs |
Kubernetes SIG Network formed an AI Gateway working group in March 2026 to make token limits, model-aware routing and provider failover standard Gateway API features, so these tools are converging on one configuration model.
Code#
A token-bucket limiter that reserves an estimate and settles the real usage, with a fallback router on top. Standard library only.
// gateway.go — token-based rate limiting with reserve-and-settle, plus fallback routing.
package main
import (
"errors"
"fmt"
"sync"
"time"
)
// bucket holds a tenant's tokens-per-minute allowance.
type bucket struct {
mu sync.Mutex
capacity float64
level float64
refill float64 // tokens per second
last time.Time
}
func newBucket(perMinute float64) *bucket {
return &bucket{capacity: perMinute, level: perMinute, refill: perMinute / 60, last: time.Now()}
}
// reserve takes an estimate up front; settle returns what was not used.
func (b *bucket) reserve(n float64) bool {
b.mu.Lock()
defer b.mu.Unlock()
now := time.Now()
b.level = min(b.capacity, b.level+now.Sub(b.last).Seconds()*b.refill)
b.last = now
if b.level < n {
return false
}
b.level -= n
return true
}
func (b *bucket) settle(reserved, used float64) {
b.mu.Lock()
defer b.mu.Unlock()
b.level = min(b.capacity, b.level+reserved-used)
}
type route struct {
name string
call func(prompt string) (out string, outTokens float64, err error)
}
var errLimited = errors.New("tenant over token limit")
// complete reserves, tries each route in order, then settles real usage.
func complete(b *bucket, routes []route, prompt string, maxTokens float64) (string, error) {
inTokens := float64(len(prompt)) / 4 // rough: four characters per token
reserved := inTokens + maxTokens
if !b.reserve(reserved) {
return "", errLimited
}
for _, r := range routes {
out, outTokens, err := r.call(prompt)
if err != nil {
fmt.Printf(" route %-8s failed: %v, falling back\n", r.name, err)
continue
}
b.settle(reserved, inTokens+outTokens)
return fmt.Sprintf("[%s] %s", r.name, out), nil
}
b.settle(reserved, 0) // nothing was generated
return "", errors.New("all routes failed")
}
func main() {
tenant := newBucket(2000) // tokens per minute
routes := []route{
{"primary", func(string) (string, float64, error) {
return "", 0, errors.New("deadline exceeded waiting for first token")
}},
{"fallback", func(string) (string, float64, error) { return "a short answer", 60, nil }},
}
for i := 1; i <= 5; i++ {
out, err := complete(tenant, routes, "Explain reserve-and-settle in one paragraph.", 500)
fmt.Printf("request %d: %q err=%v level=%.0f\n", i, out, err, tenant.level)
}
}Each request reserves about 511 tokens and uses about 71, so the bucket drains slowly. Set the fallback’s output to 500 and the fourth request is refused: that is the limiter doing its job.
Remember this#
- One gateway in front of every model: identity, token limits, routing, fallback, one usage record.
- Limits are in tokens and money. Reserve an estimate, settle the real usage.
- Fall back only before the first token, and only onto capacity that exists.
- Logical model names in application code; real model names in gateway configuration.
Try it#
- Run
gateway.go. Change the limit until the third request is refused. - Add a per-day dollar budget next to the per-minute bucket. Which should be checked first?
- Design the routing rule for a tenant whose data must stay in one region. What happens to fallback?
Check yourself#
- Why can output tokens not be limited exactly at admission time?
- Why is falling back after the first token dangerous?
- What is the one record a gateway must emit, and who reads it?