Pidoku

Platform and Operations

Intermediate 50 min Difficulty 3/5 Lesson 06 of 06

Prerequisites The Serving Layer, The Data Layer

The idea in one minute#

Compute, cluster, serving and data make AI infrastructure. The platform layer makes it something forty teams can use without talking to each other or to you. It provides a small set of paved paths — get a key, call a model, deploy an agent, see your usage — and enforces the rules on those paths so that teams do not have to remember them. Operating it means running five loops: attribute every token, observe every layer, control cost, change safely, and survive failure.

A picture#

flowchart TB
  TEAMS[":i-users: <b>Product teams</b>"] --> PORTAL[":backstage: <b>Self-service</b><br/><small>keys, quotas, catalogue, templates</small>"]
  PORTAL --> GW[":envoyproxy: <b>AI gateway</b><br/><small>the only door</small>"]
  GW --> MODELS[":vllm: <b>Models</b><br/><small>hosted and self-hosted</small>"]
  GW --> TOOLS[":modelcontextprotocol: <b>Tool gateway</b><br/><small>approved MCP servers</small>"]
  GW --> RUN[":i-bot: <b>Agent runtime</b><br/><small>+ sandboxes</small>"]
  subgraph OPS["Operations loops"]
    direction LR
    O1[":opentelemetry: <b>Observe</b>"] --> O2[":grafana: <b>SLOs and alerts</b>"]
    O1 --> O3[":i-coins: <b>Cost and showback</b>"]
    O1 --> O4[":i-scale: <b>Quality evals</b>"]
  end
  GW -.-> O1
  MODELS -.-> O1
  RUN -.-> O1
  POL[":opa: <b>Policy</b><br/><small>who may use what</small>"] --> GW
  CD[":argo: <b>GitOps delivery</b><br/><small>models, prompts, config</small>"] --> MODELS
  CD --> RUN
  class TEAMS neutral
  class PORTAL,GW queue
  class MODELS,RUN compute
  class TOOLS io
  class O1,O2,O3,O4 memory
  class POL warn
  class CD neutral

How it really works#

The paved paths#

A platform succeeds when the easy way is the governed way. Offer these and nothing looser:

A team wants toThe platform gives
Call a modelA gateway key scoped to their project, a logical model name, a quota
Use company knowledgeA retrieval API that enforces the caller’s permissions
Give an agent a toolA catalogue of approved MCP servers, reached through the gateway
Deploy an agentA runtime with checkpoints, budgets, sandboxes and tracing already wired
Self-host a modelA template: registry entry, InferencePool, autoscaling, dashboards
Know what they spentA usage and cost view, per project, updated hourly

Direct provider keys in application code are the first thing a platform removes. Every model call that bypasses the gateway is invisible to quotas, audit and cost control.

Multi-tenancy#

Tenants may be customers or internal teams; the questions are the same.

ConcernMechanism
Noisy neighboursToken and concurrency limits per tenant at the gateway; priority classes in the engine queue
Fairness on shared GPUsQuotas with borrowing in the scheduler
Data isolationTenant filters in every store; separate indexes or encryption keys for strict tenants
Cache isolationPrefix caches and semantic caches never shared across tenants that must not observe each other
Execution isolationOne sandbox per task; never a shared interpreter
Strong isolationDedicated pools, or confidential computing, for tenants that require it

Timing can leak through a shared cache: a response that arrives suspiciously fast reveals that someone else sent the same prefix. Isolating caches per tenant costs some efficiency and closes the channel. Multi-tenancy covers the serving side.

Loop 1 — attribute#

Every request carries tenant, project, user and — for agents — task and agent identity, from the gateway to the last tool call. Use the OpenTelemetry trace context to carry it. If a token cannot be attributed, it cannot be billed, limited or investigated.

Loop 2 — observe#

Four views, each answering one question:

ViewQuestionCore signals
ServiceAre users affected?TTFT, TPOT, error rate, goodput — per model and tenant tier
CapacityHow close to the limit?Queue depth, KV-cache use, occupancy, pending pods
CostWhere does the money go?Tokens and dollars per tenant, per model, per task; idle GPU-hours
QualityIs it still good?Evaluation scores by version, repair rate, user feedback

The OpenTelemetry GenAI conventions give names for model, tool and agent spans. They are still marked Development, so pin the version and keep attribute names in one shared package. The full design is in Observability Engineering.

Loop 3 — control cost#

Cost control is mostly engineering, in this order of return:

  1. Cache the prefix — largest saving, no behaviour change.
  2. Route easy calls to small models — several-fold on those calls.
  3. Cap output and context — max_tokens, compaction, trimming tool results.
  4. Batch what is not urgent — batch APIs and off-peak GPU time are far cheaper.
  5. Raise occupancy — the same fleet at 60% instead of 30% halves unit cost.
  6. Budgets with teeth — per-project daily limits that stop traffic, not just alert.

Publish cost per unit of work — per ticket resolved, per pull request — not per token. That is the number a product owner can act on.

Loop 4 — change safely#

In an AI platform, the things that change behaviour are not only code:

model version      prompt templates      tool descriptions      routing rules
retrieval config   guardrail policies    engine version         quantization

All of them live in version control, go through review, run the evaluation gate, and roll out by canary. “Prompt as configuration edited in a console” is how regressions reach production with no record of who changed what.

Loop 5 — survive failure#

FailurePreparation
A provider outageA second route with tested prompts and reserved capacity
A regional outageStateless serving in two regions; state replicated; weights pre-staged
A bad model or prompt rolloutCanary plus one-step rollback
A GPU or node failureN+1 replicas; health-based tainting; automatic rescheduling
A runaway agentBudgets, a kill switch per task and per tenant
A security incidentRevocable per-agent credentials; audit trail; a practised response

Decide which workloads are worth a second region. Interactive serving usually is; a batch evaluation queue is not. Multi-region and DR covers the mechanics.

The platform team’s interface#

A platform is a product. Its users judge it on four things:

  • Time to first call for a new team: minutes, self-service.
  • Clear limits: published quotas, a predictable way to raise them.
  • Honest status: SLOs per model, visible to tenants.
  • A deprecation policy: models are retired on a schedule, with notice and a migration path.

Remember this#

  • The platform’s job is paved paths: the governed way is the easy way.
  • One door — the gateway — for models and tools. No direct provider keys.
  • Five loops: attribute, observe, control cost, change safely, survive failure.
  • Everything that changes behaviour is versioned and passes the evaluation gate.
  • Report cost per unit of work.

Try it#

  1. List the behaviour-changing artifacts in a system you know. Which are not in version control?
  2. Design the tenant limits for a platform with a free tier, a paid tier and internal batch.
  3. Choose one failure from the table and write the runbook’s first five steps.

Check yourself#

  1. Why does a direct provider key in an application undermine the platform?
  2. How can a shared prefix cache leak information between tenants?
  3. Which six cost levers exist, and which has the highest return for the least risk?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom