Pidoku

AI Infrastructure for an Agentic Platform

Expert 1h 30m Difficulty 5/5 Lesson 01 of 03

Prerequisites all earlier topics in this course

The idea in one minute#

This lesson designs, from an empty page, the infrastructure for a company-wide agentic AI platform: the shared system on which product teams run agents that write code, resolve support tickets and research questions. It is worked in the six steps of the design method. The result is seven planes — edge, agent, model, tool, execution, data, and control — each built from components covered earlier, sized with arithmetic, and checked against failure and attack.

The shape of the answer is worth stating first, because it surprises people: an agent platform is mostly not GPUs. It is a durable workflow system, a sandbox fleet and a policy layer, with model serving as one plane among seven.

Step 1 — requirements#

The problem. A 5,000-engineer company wants one platform for agentic workloads instead of a dozen team-built ones with separate keys, no audit trail and unpredictable bills.

RequirementDecision
WorkloadsCoding agents (long tasks in a repository), support agents (ticket resolution with tools), research agents (search and synthesis)
Users5,000 engineers, 400 support staff; 40 product teams building on the platform
Scale at peak1,500 concurrent agent tasks; 12,000 tasks per hour
Task shape5 to 100 model calls; seconds to hours; may wait on a human for a day
LatencyInteractive steps: first token under 1.5 s. Background tasks: throughput over latency
QualityPer-workload evaluation suites; no release without passing
DurabilityNo task loses progress to a crash or a deploy
DataSource code and customer tickets must stay inside the company’s cloud boundary for the self-hosted path; hosted frontier models allowed for code under a zero-retention agreement
SafetyNo irreversible external action without approval; every action attributable to a user
TenancyTeams isolated from each other in budget, data and execution
Availability99.9% for the platform; graceful degradation when a model provider fails

Step 2 — numbers#

concurrent tasks at peak                         1,500
model calls in flight (a task is in a model
call ~35% of the time; the rest is tools,
sandboxes and waiting)                             ~525
average call: input 35,000 tokens, 85% cached;
              output 500 tokens; ~12 s

calls per second            525 ÷ 12 s           ≈ 44
input tokens per second     44 × 35,000          ≈ 1.5 M   (of which ~0.23 M uncached)
output tokens per second    44 × 500             ≈ 22,000

Two conclusions drive the whole design:

  1. Input outweighs output seventy to one. The serving plane is a prefill-and-cache problem. Prefix-cache hit rate is the most valuable number on the platform.
  2. Two-thirds of a task’s life is not in a model call. The agent runtime, sandboxes and tools are where most of the time goes; they need as much design as the GPUs.

The model split, decided by evaluation per call site:

frontier model, hosted      planning and hard reasoning steps        ~25% of calls
large open model, self-hosted   routine agent steps, tool use        ~55% of calls
small models, self-hosted   routing, classification, guardrails,
                            embeddings, reranking, compaction        ~20% of calls + all embeddings

The program at the end of the lesson turns this into GPU counts, sandbox capacity and cost.

Step 3 — architecture#

flowchart TB
  subgraph EDGE["1 Edge plane"]
    direction LR
    CL[":i-users: <b>IDE, web, API, chat</b>"] --> IDP[":keycloak: <b>SSO</b>"]
    IDP --> GW[":envoyproxy: <b>AI gateway</b><br/><small>identity, budgets, routing</small>"]
  end
  subgraph AGENT["2 Agent plane"]
    direction LR
    API[":i-server: <b>Task API</b><br/><small>create, stream, approve</small>"] --> WF[":temporal: <b>Durable workflows</b><br/><small>one per task</small>"]
    WF --> RT[":i-bot: <b>Agent workers</b><br/><small>the loop, context builder</small>"]
    RT --> POL[":opa: <b>Policy layer</b>"]
  end
  subgraph MODEL["3 Model plane"]
    direction LR
    MR[":envoyproxy: <b>Model routes</b>"] --> HOST[":anthropic: <b>Hosted frontier</b>"]
    MR --> EPP[":llm-d: <b>Endpoint picker</b>"]
    EPP --> VL[":vllm: <b>Large model pool</b>"]
    EPP --> SM[":vllm: <b>Small model pool</b>"]
  end
  subgraph TOOL["4 Tool plane"]
    direction LR
    TG[":agentgateway: <b>Tool gateway</b>"] --> MCPS[":modelcontextprotocol: <b>MCP servers</b><br/><small>repos, tickets, search, docs</small>"]
    TG --> A2[":a2a: <b>Partner agents</b>"]
  end
  subgraph EXEC["5 Execution plane"]
    direction LR
    SC[":kubernetes: <b>Sandbox controller</b><br/><small>warm pool</small>"] --> SB[":gvisor: <b>Sandboxes</b>"]
    SB --> EG[":cilium: <b>Egress allowlist</b>"]
  end
  subgraph DATA["6 Data plane"]
    direction TB
    PG[(":postgresql: <b>Tasks, memory</b>")]
    RD[(":redis: <b>Sessions, limits</b>")]
    VEC[(":qdrant: <b>Search index</b>")]
    OBJ[(":minio: <b>Artifacts, weights</b>")]
  end
  subgraph CTRL["7 Control plane"]
    direction LR
    OT[":opentelemetry: <b>Telemetry</b>"] --> GF[":grafana: <b>SLOs, cost</b>"]
    EV[":i-scale: <b>Evaluation</b>"]
    CD[":argo: <b>GitOps</b>"]
    KQ[":kueue: <b>Quotas</b>"]
  end
  EDGE --> AGENT
  AGENT --> MODEL
  AGENT --> TOOL
  AGENT --> EXEC
  AGENT --> DATA
  MODEL --> GPU[(":nvidia: <b>GPU node pools</b>")]
  CTRL -.-> AGENT
  CTRL -.-> MODEL
  class CL neutral
  class IDP memory
  class GW,MR,TG,POL,EPP queue
  class API,WF,RT compute
  class HOST,MCPS,A2 io
  class VL,SM compute
  class SC neutral
  class SB,EG warn
  class PG,RD,VEC,OBJ,GPU memory
  class OT,GF,EV,CD,KQ neutral

Plane 1 — edge#

Users reach the platform from an IDE extension, a web console, an API and chat. Single sign-on issues a user token; the AI gateway authenticates it, attributes the request to tenant, project and user, applies token-per-minute limits and daily budgets, and emits the usage event. Product teams receive gateway keys scoped to a project. There are no provider keys anywhere outside the gateway. (Model Access and Gateways)

Plane 2 — agent#

  • Task API. Create a task, stream its events, send an approval, cancel. Each task has an ID, an owner and budgets for steps, tokens, time and spend.
  • Durable workflows. One workflow per task. Each model call and tool call is a recorded activity, so a task resumes on any worker after a crash or deploy, and can wait a day for a human without holding a process. (Orchestration and Durable Execution)
  • Agent workers. Stateless processes that run the loop: build the context, call the model through the gateway, propose an action. They scale on the task queue’s depth. Context management — tool-result clearing, compaction, sub-agents — lives here. (Context Engineering and Memory)
  • Policy layer. Every proposed tool call is checked in code: is the tool allowed for this task, are the arguments within bounds, has the session read untrusted content, is approval required. (A Secure Reference Architecture)

Teams build agents on this plane by supplying a system prompt, a tool set, an evaluation suite and budgets — not by running their own loop.

Plane 3 — model#

All model calls go back through the gateway to model routes named by purpose: reason-large, agent-default, fast-small, embed, rerank.

  • reason-large → a hosted frontier model, with a second provider as fallback.
  • agent-default → the self-hosted large open-weights model, with the hosted route as overflow when the pool is saturated.
  • fast-small, embed, rerank → small self-hosted models on shared GPUs.

Self-hosted pools are InferencePools. The endpoint picker routes on queue depth, KV-cache utilisation and prefix affinity: every step of a task goes back to the replica that holds that task’s cached prefix. With 85% of input repeating between steps, prefix-aware routing is worth more than any other serving optimisation here. Pools autoscale on queue wait and KV pressure, with warm spares for the morning ramp. (The Serving Layer)

Plane 4 — tool#

Agents reach tools through a tool gateway that only knows approved MCP servers: source control, the ticket system, internal search, documentation, the CI system. Since MCP’s 2026-07-28 revision these servers are stateless HTTP services, and the method and tool name travel in headers, so the gateway authorises and meters per tool without parsing bodies. Each call carries a short-lived token issued for that user and that server. Third-party agents are reached over A2A through the same gateway, and their output is treated as untrusted input. (Tools and Protocols)

Plane 5 — execution#

Coding and data-analysis agents run commands in a sandbox: one per task, gVisor-isolated, on a dedicated CPU node pool, with only that task’s repository mounted and no credentials inside. Egress is denied except to an allowlist — the internal package mirror and source host — through a proxy that injects credentials. Sandboxes come from a warm pool, pause to a snapshot while the task waits, and are destroyed at the end. (Sandboxes and Tool Execution)

Plane 6 — data#

StoreHolds
PostgresWorkflow history, task records, memory, approvals, audit trail, tenants and keys
RedisRate-limit counters, session buffers, event streams for live clients
Search indexCode and document embeddings with keyword index and ACL filters
Object storageTask artifacts, sandbox snapshots, model weights (as OCI artifacts), ingested documents
Stream + column storeUsage events, traces, evaluation results

(The Data Layer)

Plane 7 — control#

GitOps delivers everything that changes behaviour — prompts, tool definitions, policies, routing, model versions — through review, the evaluation gate and a canary. Kueue enforces GPU quotas between serving, batch evaluation and experiments. OpenTelemetry carries one trace per task across all planes; dashboards show SLOs, capacity, cost per task and quality by version. (Platform and Operations)

The life of one task#

sequenceDiagram
  participant U as Engineer
  participant G as AI gateway
  participant W as Workflow + agent worker
  participant M as Model plane
  participant P as Policy layer
  participant S as Sandbox
  participant T as Tool gateway (MCP)
  U->>G: "Fix the failing test in service-x"
  G->>W: create task (user, budgets)
  W->>S: claim sandbox, clone repo
  loop until done or budget reached
    W->>M: model call (prefix cached on same replica)
    M-->>W: proposed action
    W->>P: check action
    alt run tests / edit files
      P-->>W: allow
      W->>S: execute
      S-->>W: output (trimmed)
    else open pull request
      P-->>W: needs approval
      W->>U: show exact diff and target
      U-->>W: approve
      W->>T: create PR (user-scoped token)
    end
    W->>W: checkpoint
  end
  W->>U: result + trace link

Step 4 — failure#

FailureWhat happens
Agent worker crashes or is redeployedThe workflow resumes on another worker from its history; no step repeats its effect because tool calls carry idempotency keys
Hosted model provider slow or downThe breaker opens on slow first tokens; reason-large falls back to the second provider; tasks continue, somewhat slower
Self-hosted pool saturatedThe queue grows; the endpoint picker sheds background traffic first; agent-default overflows to the hosted route within budget
A GPU failsIts replica is tainted and rescheduled; N+1 capacity absorbs it; tasks pinned to it lose their cached prefix and pay one slow step
Sandbox node lostTasks restore from the last snapshot on another node
A tool is downThe activity retries with backoff; the agent is told the tool is unavailable and may proceed another way
An agent loopsThe no-progress detector or the step budget ends the task in a saved state with a clear reason
A bad prompt or model releaseThe canary shows a quality drop; rollback is a routing change
Postgres primary failsA replica is promoted; workflows pause for seconds and continue
Region lostInteractive traffic fails over to a second region with replicated state and pre-staged weights; long background tasks resume there from history

Step 5 — cost#

The program below produces this breakdown for the stated load:

component                              per day
─────────                              ───────
hosted frontier model (25% of calls)   the largest line; sensitive to cache hit rate
self-hosted large model GPUs           fixed by replica count; cheap per token when busy
self-hosted small model GPUs           small
sandboxes (CPU)                        modest; pause-on-idle cuts it several-fold
platform services, storage, telemetry  a few percent

Levers, in order of effect on this design:

  1. Prefix-cache hit rate. From 85% to 60% roughly doubles uncached input. Keep prompts append-only; route by prefix; do not put timestamps in system prompts.
  2. The model split. Each call site moved from the frontier route to the self-hosted route — where evaluation shows quality holds — moves cost from per-token to already-paid GPUs.
  3. Context discipline. Clearing tool results and compacting cuts average input per call.
  4. Batch what can wait. Nightly evaluation and background research run on spare capacity.
  5. Pause idle sandboxes.

Report cost per completed task by workload. A coding task that costs $1.40 and saves an engineer forty minutes is easy to justify; the same task at $14 because caching broke is a bug to find within the hour.

Step 6 — security#

Apply the worksheet from Trust Boundaries:

ItemIn this design
Untrusted sourcesTicket text from customers; web pages fetched by research agents; file contents in repositories; tool results; partner agents
SinksSandbox commands; pull requests; ticket replies; outbound HTTP; memory writes
TrifectaSupport agents hold customer data, read customer-written text and can send replies → replies to external addresses go through approval or a templated, validated path. Research agents read the web and have no access to private data. Coding agents read untrusted repository content but have no network egress beyond the allowlist
IdentityEvery tool call uses a token issued for the requesting user, scoped to the task’s resources, valid for minutes
IsolationOne sandbox per task; tenants never share a sandbox, a prefix cache or a semantic cache
Supply chainModels from the internal registry, signed; MCP servers from an approved, pinned catalogue
Blast radius of a hijacked coding agentIt can damage its own sandbox and propose a pull request to one repository, which a human reviews. It cannot reach other repositories, secrets or the internet
DetectionAudit records with context provenance; alerts on first-time tool use after untrusted input, on egress denials, on approval-rate anomalies

The full security treatment of this same platform is in Securing an Agent Platform.

Rollout#

No platform of this size is built in one step. An order in which each stage is useful alone:

  1. Gateway in front of hosted models: keys, budgets, usage events.
  2. Tracing and evaluation harness; first workload’s evaluation suite.
  3. Agent plane with durable workflows and the policy layer, for one team, hosted models only.
  4. Sandboxes with egress control; coding agents become possible.
  5. Tool gateway and the approved MCP catalogue.
  6. Self-hosted model pool for agent-default, once usage data proves steady volume.
  7. Quotas, showback and self-service as teams multiply.
  8. Second region for the interactive path.

Code#

The capacity and cost model for this design.

Go
// platform.go — capacity and daily cost of the agentic platform design.
package main

import (
	"fmt"
	"math"
)

func main() {
	const (
		concurrentTasks = 1500.0
		inModelCall     = 0.35 // share of a task's life spent in a model call
		callSeconds     = 12.0
		tokIn           = 35000.0
		tokOut          = 500.0
		cacheHit        = 0.85

		shareFrontier = 0.25 // hosted
		shareLarge    = 0.55 // self-hosted large model
		// the remaining 20% are small-model calls

		priceIn, priceCached, priceOut = 3.00, 0.30, 15.00 // $ per million tokens, hosted

		largeGoodput   = 1300.0 // output tokens/s per replica at the SLO, measured with these prompts
		gpusPerReplica = 2.0
		smallGPUs      = 6.0
		gpuHour        = 3.50
		occupancy      = 0.6

		sandboxActive = 0.15 // share of tasks whose sandbox is running rather than paused
		sandboxHour   = 0.05 // $ per active sandbox-hour (1 vCPU, 2 GB)
		peakHours     = 10.0 // peak-equivalent hours per day
	)

	callsPerSec := concurrentTasks * inModelCall / callSeconds
	fmt.Printf("model calls in flight %.0f   calls/s %.1f\n", concurrentTasks*inModelCall, callsPerSec)
	fmt.Printf("input tokens/s %.2f M (uncached %.2f M)   output tokens/s %.0f\n\n",
		callsPerSec*tokIn/1e6, callsPerSec*tokIn*(1-cacheHit)/1e6, callsPerSec*tokOut)

	// hosted frontier route
	perCall := (tokIn*(1-cacheHit)*priceIn + tokIn*cacheHit*priceCached + tokOut*priceOut) / 1e6
	frontierDay := callsPerSec * shareFrontier * perCall * 3600 * peakHours

	// self-hosted large model
	largeOut := callsPerSec * shareLarge * tokOut
	replicas := math.Ceil(largeOut/largeGoodput/occupancy) + 1 // +1 spare
	largeGPUs := replicas * gpusPerReplica
	gpuDay := (largeGPUs + smallGPUs) * gpuHour * 24

	sandboxDay := concurrentTasks * sandboxActive * sandboxHour * peakHours

	fmt.Printf("hosted frontier:   $%.4f per call   $%8.0f / day\n", perCall, frontierDay)
	fmt.Printf("large model pool:  %.0f replicas, %.0f GPUs\n", replicas, largeGPUs)
	fmt.Printf("all GPUs:          %.0f GPUs          $%8.0f / day\n", largeGPUs+smallGPUs, gpuDay)
	fmt.Printf("sandboxes:         %.0f active        $%8.0f / day\n", concurrentTasks*sandboxActive, sandboxDay)

	total := frontierDay + gpuDay + sandboxDay
	tasksPerDay := 12000.0 * peakHours
	fmt.Printf("\ntotal ≈ $%.0f / day   → $%.2f per task over %.0f tasks\n", total, total/tasksPerDay, tasksPerDay)

	// the sensitivity that matters most
	low := (tokIn*0.4*priceIn + tokIn*0.6*priceCached + tokOut*priceOut) / 1e6
	fmt.Printf("\nif the cache hit rate falls to 60%%, a frontier call costs $%.4f (×%.1f)\n", low, low/perCall)
}

Run it, then change one thing at a time. Moving shareFrontier from 0.25 to 0.10 and cacheHit from 0.85 to 0.60 are the two experiments that teach the most.

Remember this#

  • An agent platform is seven planes: edge, agent, model, tool, execution, data, control.
  • Input tokens dominate; prefix caching and prefix-aware routing dominate cost and latency.
  • Most of a task’s time is outside the model; durable workflows and sandboxes carry the weight.
  • Every tool call crosses a policy check with the user’s scoped, short-lived authority.
  • Build it gateway-first, one useful stage at a time.

Try it#

  1. Run platform.go with your organisation’s size. Which line dominates?
  2. Add a fourth workload: a browser-using agent. Which planes change, and what does the trifecta analysis say?
  3. Remove the self-hosted large model and send everything to hosted APIs. At what daily task volume does self-hosting pay back?

Check yourself#

  1. Why is prefix affinity the most valuable routing signal in this design?
  2. Which plane guarantees that a deploy does not lose an hour-long task?
  3. What limits the damage of a hijacked coding agent here?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom