The idea in one minute#
Compute, cluster, serving and data make AI infrastructure. The platform layer makes it something forty teams can use without talking to each other or to you. It provides a small set of paved paths — get a key, call a model, deploy an agent, see your usage — and enforces the rules on those paths so that teams do not have to remember them. Operating it means running five loops: attribute every token, observe every layer, control cost, change safely, and survive failure.
A picture#
flowchart TB
TEAMS[":i-users: <b>Product teams</b>"] --> PORTAL[":backstage: <b>Self-service</b><br/><small>keys, quotas, catalogue, templates</small>"]
PORTAL --> GW[":envoyproxy: <b>AI gateway</b><br/><small>the only door</small>"]
GW --> MODELS[":vllm: <b>Models</b><br/><small>hosted and self-hosted</small>"]
GW --> TOOLS[":modelcontextprotocol: <b>Tool gateway</b><br/><small>approved MCP servers</small>"]
GW --> RUN[":i-bot: <b>Agent runtime</b><br/><small>+ sandboxes</small>"]
subgraph OPS["Operations loops"]
direction LR
O1[":opentelemetry: <b>Observe</b>"] --> O2[":grafana: <b>SLOs and alerts</b>"]
O1 --> O3[":i-coins: <b>Cost and showback</b>"]
O1 --> O4[":i-scale: <b>Quality evals</b>"]
end
GW -.-> O1
MODELS -.-> O1
RUN -.-> O1
POL[":opa: <b>Policy</b><br/><small>who may use what</small>"] --> GW
CD[":argo: <b>GitOps delivery</b><br/><small>models, prompts, config</small>"] --> MODELS
CD --> RUN
class TEAMS neutral
class PORTAL,GW queue
class MODELS,RUN compute
class TOOLS io
class O1,O2,O3,O4 memory
class POL warn
class CD neutralHow it really works#
The paved paths#
A platform succeeds when the easy way is the governed way. Offer these and nothing looser:
| A team wants to | The platform gives |
|---|---|
| Call a model | A gateway key scoped to their project, a logical model name, a quota |
| Use company knowledge | A retrieval API that enforces the caller’s permissions |
| Give an agent a tool | A catalogue of approved MCP servers, reached through the gateway |
| Deploy an agent | A runtime with checkpoints, budgets, sandboxes and tracing already wired |
| Self-host a model | A template: registry entry, InferencePool, autoscaling, dashboards |
| Know what they spent | A usage and cost view, per project, updated hourly |
Direct provider keys in application code are the first thing a platform removes. Every model call that bypasses the gateway is invisible to quotas, audit and cost control.
Multi-tenancy#
Tenants may be customers or internal teams; the questions are the same.
| Concern | Mechanism |
|---|---|
| Noisy neighbours | Token and concurrency limits per tenant at the gateway; priority classes in the engine queue |
| Fairness on shared GPUs | Quotas with borrowing in the scheduler |
| Data isolation | Tenant filters in every store; separate indexes or encryption keys for strict tenants |
| Cache isolation | Prefix caches and semantic caches never shared across tenants that must not observe each other |
| Execution isolation | One sandbox per task; never a shared interpreter |
| Strong isolation | Dedicated pools, or confidential computing, for tenants that require it |
Timing can leak through a shared cache: a response that arrives suspiciously fast reveals that someone else sent the same prefix. Isolating caches per tenant costs some efficiency and closes the channel. Multi-tenancy covers the serving side.
Loop 1 — attribute#
Every request carries tenant, project, user and — for agents — task and agent identity, from the gateway to the last tool call. Use the OpenTelemetry trace context to carry it. If a token cannot be attributed, it cannot be billed, limited or investigated.
Loop 2 — observe#
Four views, each answering one question:
| View | Question | Core signals |
|---|---|---|
| Service | Are users affected? | TTFT, TPOT, error rate, goodput — per model and tenant tier |
| Capacity | How close to the limit? | Queue depth, KV-cache use, occupancy, pending pods |
| Cost | Where does the money go? | Tokens and dollars per tenant, per model, per task; idle GPU-hours |
| Quality | Is it still good? | Evaluation scores by version, repair rate, user feedback |
The OpenTelemetry GenAI conventions give names for model, tool and agent spans. They are still marked Development, so pin the version and keep attribute names in one shared package. The full design is in Observability Engineering.
Loop 3 — control cost#
Cost control is mostly engineering, in this order of return:
- Cache the prefix — largest saving, no behaviour change.
- Route easy calls to small models — several-fold on those calls.
- Cap output and context —
max_tokens, compaction, trimming tool results. - Batch what is not urgent — batch APIs and off-peak GPU time are far cheaper.
- Raise occupancy — the same fleet at 60% instead of 30% halves unit cost.
- Budgets with teeth — per-project daily limits that stop traffic, not just alert.
Publish cost per unit of work — per ticket resolved, per pull request — not per token. That is the number a product owner can act on.
Loop 4 — change safely#
In an AI platform, the things that change behaviour are not only code:
model version prompt templates tool descriptions routing rules
retrieval config guardrail policies engine version quantizationAll of them live in version control, go through review, run the evaluation gate, and roll out by canary. “Prompt as configuration edited in a console” is how regressions reach production with no record of who changed what.
Loop 5 — survive failure#
| Failure | Preparation |
|---|---|
| A provider outage | A second route with tested prompts and reserved capacity |
| A regional outage | Stateless serving in two regions; state replicated; weights pre-staged |
| A bad model or prompt rollout | Canary plus one-step rollback |
| A GPU or node failure | N+1 replicas; health-based tainting; automatic rescheduling |
| A runaway agent | Budgets, a kill switch per task and per tenant |
| A security incident | Revocable per-agent credentials; audit trail; a practised response |
Decide which workloads are worth a second region. Interactive serving usually is; a batch evaluation queue is not. Multi-region and DR covers the mechanics.
The platform team’s interface#
A platform is a product. Its users judge it on four things:
- Time to first call for a new team: minutes, self-service.
- Clear limits: published quotas, a predictable way to raise them.
- Honest status: SLOs per model, visible to tenants.
- A deprecation policy: models are retired on a schedule, with notice and a migration path.
Remember this#
- The platform’s job is paved paths: the governed way is the easy way.
- One door — the gateway — for models and tools. No direct provider keys.
- Five loops: attribute, observe, control cost, change safely, survive failure.
- Everything that changes behaviour is versioned and passes the evaluation gate.
- Report cost per unit of work.
Try it#
- List the behaviour-changing artifacts in a system you know. Which are not in version control?
- Design the tenant limits for a platform with a free tier, a paid tier and internal batch.
- Choose one failure from the table and write the runbook’s first five steps.
Check yourself#
- Why does a direct provider key in an application undermine the platform?
- How can a shared prefix cache leak information between tenants?
- Which six cost levers exist, and which has the highest return for the least risk?