Pidoku

Model Access and Gateways

Basic 45 min Difficulty 2/5 Lesson 01 of 06

Prerequisites The Design Method

The idea in one minute#

Applications should never call a model directly. They call an AI gateway, and the gateway calls models. That one indirection is where every cross-cutting concern lives: who is calling, how many tokens they may spend, which model serves this request, what happens when that model is slow, and what gets logged. Without it, each of those is reimplemented — differently — in every service.

An AI gateway differs from an ordinary API gateway in three ways: it meters tokens, it routes on model state, and it handles streams that last half a minute.

A picture#

flowchart LR
  APP[":i-code: <b>Applications</b><br/><small>one API, one key each</small>"] --> GW
  subgraph GW["AI gateway"]
    direction TB
    A1[":i-fingerprint: Authenticate, attribute"] --> A2[":i-coins: Token limits and budgets"]
    A2 --> A3[":i-route: Route by model, cost, health"]
    A3 --> A4[":i-shield-check: Guardrail hooks"]
  end
  GW --> P1[":anthropic: <b>Hosted API A</b>"]
  GW --> P2[":googlegemini: <b>Hosted API B</b>"]
  GW --> P3[":vllm: <b>Your own pool</b><br/><small>InferencePool on Kubernetes</small>"]
  P1 -.->|"slow or failing"| P2
  GW -.-> T[":opentelemetry: <b>Usage events</b><br/><small>tokens, latency, tenant</small>"]
  class APP neutral
  class A1,A2,A3 queue
  class A4 warn
  class P1,P2 io
  class P3 compute
  class T memory

How it really works#

Hosted, managed or self-hosted#

The first decision is where the model runs. It is a ladder; climb only when a reason forces you.

RungYou operateChoose it whenYou give up
Hosted APINothingStarting out; you need frontier quality; volume is modest or spikyControl of latency, data location, price
Dedicated capacity from a providerNothing, but you reserve throughputLatency must be predictable; volume is steadyFlexibility; you pay for idle
Managed open-weights endpointConfigurationYou need a specific open model or fine-tune without running GPUsSome efficiency and tuning freedom
Self-hosted on your GPUsEverythingData cannot leave; volume is large and steady; you need custom models or the lowest unit costSimplicity. This is a platform team

Most real systems use two rungs at once: a frontier API for the hard calls and a self-hosted or small model for the many easy ones. The gateway is what makes that mixture invisible to application code. AI Infrastructure From Scratch covers the bottom rung.

What the gateway does#

Identity and attribution. Applications hold a gateway key, never a provider key. Each request is attributed to a tenant, application and user, and that attribution rides on every usage event. Provider credentials exist in exactly one place.

Token-based limits. Limits are in tokens per minute and dollars per day, per tenant. There is a wrinkle: output tokens are not known when the request arrives. Gateways handle this by reserving an estimate (input tokens plus max_tokens) and settling the difference when the response finishes.

Routing. A request names a logical model — “fast”, “smart”, “code” — and the gateway maps it to a real one. Routing can consider price, health, the tenant’s data-residency rule and, for your own pool, live engine state: the Gateway API Inference Extension routes to the replica with the shortest queue and the warmest cache rather than round-robin.

Fallback. When a route errors or exceeds its deadline for first token, the gateway retries on another. Two rules keep fallback safe: fall back only before the first token has been sent to the client, and send fallback traffic to a route with spare capacity — otherwise one provider’s outage becomes two.

Caching. Provider-side prompt caching needs stable prefixes, and the gateway is where you can enforce prompt layout and measure the cache hit rate. Exact-match response caching is safe for deterministic calls; semantic caching is covered, with its risks, in State, Memory and Caching.

Guardrail hooks. The gateway is the natural place to call input and output checks, because every request already passes through it. What those checks can and cannot do is the subject of Runtime Guardrails.

One record per request. Tenant, model, route, input, cached and output tokens, TTFT, total time, finish reason, cost. This single wide event is the source of truth for billing, capacity planning and incident review.

Model routing as a cost lever#

Not every call needs the largest model. A router that sends easy calls to a small model cuts cost several-fold, and the design question is how it decides:

StrategyHowTrade-off
By taskThe application says which tier each call site needsSimple and predictable; needs discipline
CascadeTry the small model; escalate when a check failsSaves most; adds latency on escalation
Learned routerA classifier predicts which model is sufficientBest savings; another model to evaluate

Whatever the strategy, changes to routing are behaviour changes and go through the same evaluation gate as a model upgrade.

What fills the box in 2026#

ToolNotes
Envoy AI GatewayBuilt on Envoy Gateway; 1.0 in June 2026. Token-based rate limiting, provider failover, MCP routing, OpenTelemetry GenAI tracing
agentgatewayRust data plane aimed at agent traffic: LLM, MCP and A2A in one proxy; integrates with InferencePool routing
LiteLLMPython proxy with the widest provider coverage; quick to adopt
Kong, cloud provider gatewaysSensible when you already run them for other APIs

Kubernetes SIG Network formed an AI Gateway working group in March 2026 to make token limits, model-aware routing and provider failover standard Gateway API features, so these tools are converging on one configuration model.

Code#

A token-bucket limiter that reserves an estimate and settles the real usage, with a fallback router on top. Standard library only.

Go
// gateway.go — token-based rate limiting with reserve-and-settle, plus fallback routing.
package main

import (
	"errors"
	"fmt"
	"sync"
	"time"
)

// bucket holds a tenant's tokens-per-minute allowance.
type bucket struct {
	mu       sync.Mutex
	capacity float64
	level    float64
	refill   float64 // tokens per second
	last     time.Time
}

func newBucket(perMinute float64) *bucket {
	return &bucket{capacity: perMinute, level: perMinute, refill: perMinute / 60, last: time.Now()}
}

// reserve takes an estimate up front; settle returns what was not used.
func (b *bucket) reserve(n float64) bool {
	b.mu.Lock()
	defer b.mu.Unlock()
	now := time.Now()
	b.level = min(b.capacity, b.level+now.Sub(b.last).Seconds()*b.refill)
	b.last = now
	if b.level < n {
		return false
	}
	b.level -= n
	return true
}

func (b *bucket) settle(reserved, used float64) {
	b.mu.Lock()
	defer b.mu.Unlock()
	b.level = min(b.capacity, b.level+reserved-used)
}

type route struct {
	name string
	call func(prompt string) (out string, outTokens float64, err error)
}

var errLimited = errors.New("tenant over token limit")

// complete reserves, tries each route in order, then settles real usage.
func complete(b *bucket, routes []route, prompt string, maxTokens float64) (string, error) {
	inTokens := float64(len(prompt)) / 4 // rough: four characters per token
	reserved := inTokens + maxTokens
	if !b.reserve(reserved) {
		return "", errLimited
	}
	for _, r := range routes {
		out, outTokens, err := r.call(prompt)
		if err != nil {
			fmt.Printf("  route %-8s failed: %v, falling back\n", r.name, err)
			continue
		}
		b.settle(reserved, inTokens+outTokens)
		return fmt.Sprintf("[%s] %s", r.name, out), nil
	}
	b.settle(reserved, 0) // nothing was generated
	return "", errors.New("all routes failed")
}

func main() {
	tenant := newBucket(2000) // tokens per minute
	routes := []route{
		{"primary", func(string) (string, float64, error) {
			return "", 0, errors.New("deadline exceeded waiting for first token")
		}},
		{"fallback", func(string) (string, float64, error) { return "a short answer", 60, nil }},
	}
	for i := 1; i <= 5; i++ {
		out, err := complete(tenant, routes, "Explain reserve-and-settle in one paragraph.", 500)
		fmt.Printf("request %d: %q err=%v level=%.0f\n", i, out, err, tenant.level)
	}
}

Each request reserves about 511 tokens and uses about 71, so the bucket drains slowly. Set the fallback’s output to 500 and the fourth request is refused: that is the limiter doing its job.

Remember this#

  • One gateway in front of every model: identity, token limits, routing, fallback, one usage record.
  • Limits are in tokens and money. Reserve an estimate, settle the real usage.
  • Fall back only before the first token, and only onto capacity that exists.
  • Logical model names in application code; real model names in gateway configuration.

Try it#

  1. Run gateway.go. Change the limit until the third request is refused.
  2. Add a per-day dollar budget next to the per-minute bucket. Which should be checked first?
  3. Design the routing rule for a tenant whose data must stay in one region. What happens to fallback?

Check yourself#

  1. Why can output tokens not be limited exactly at admission time?
  2. Why is falling back after the first token dangerous?
  3. What is the one record a gateway must emit, and who reads it?

Sources#

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom