The idea in one minute#
AI infrastructure is five layers stacked on top of each other: compute, cluster, serving, data and platform. You can rent the stack at any height — a model API rents all five; bare GPUs rent only the first. The design decision is not “should we build AI infrastructure?” but “at which layer do we start owning it?”, and the honest answer for most teams is higher up than they first think. Owning a layer buys control and unit cost; it costs people, time and risk.
Build the stack in the order that delivers value: gateway first, then data, then serving, and compute last.
A picture#
flowchart TB
subgraph L5["5 Platform: what teams see"]
direction TB
PL1[":envoyproxy: <b>Gateway</b><br/><small>keys, quotas, routing</small>"]
PL2[":grafana: <b>Observability</b><br/><small>usage, cost, quality</small>"]
PL3[":i-list-checks: <b>Model catalogue</b><br/><small>what is approved</small>"]
end
subgraph L4["4 Data"]
direction TB
DA1[(":i-archive: <b>Object storage</b><br/><small>weights, documents</small>")]
DA2[(":qdrant: <b>Vector index</b>")]
DA3[(":postgresql: <b>State</b>")]
end
subgraph L3["3 Serving"]
direction TB
SV1[":vllm: <b>Inference engines</b>"]
SV2[":llm-d: <b>Routing and scaling</b>"]
end
subgraph L2["2 Cluster"]
direction TB
CL1[":kubernetes: <b>Scheduler</b><br/><small>DRA, queues, quotas</small>"]
CL2[":nvidia: <b>GPU Operator</b><br/><small>drivers, device plugins</small>"]
end
subgraph L1["1 Compute"]
direction TB
CO1[":nvidia: <b>GPU nodes</b>"]
CO2[":i-network: <b>Network</b>"]
CO3[":i-hard-drive: <b>Fast local storage</b>"]
end
L5 --> L3
L5 --> L4
L3 --> L2 --> L1
L4 --> L2
class PL1,PL2,PL3 queue
class DA1,DA2,DA3 memory
class SV1,SV2 compute
class CL1,CL2 neutral
class CO1,CO2,CO3 ioHow it really works#
What each layer is responsible for#
| Layer | Responsibility | Fails as |
|---|---|---|
| Compute | Accelerators with enough memory and bandwidth; network between them; storage that loads weights quickly | No capacity; a failed GPU; slow model loads |
| Cluster | Turn machines into a pool: schedule work onto devices, enforce quotas, queue what does not fit | Pending pods; one team starving another; fragmentation |
| Serving | Turn weights into an endpoint: batching, caching, routing, autoscaling, rollouts | High TTFT; cold starts; a bad model version |
| Data | Hold everything that is not compute: weights, corpora, vectors, sessions, events | Stale indexes; lost state; permission leaks |
| Platform | Make it usable and governable by many teams: one entry point, attribution, cost, policy | Ungoverned spend; no audit trail; every team reinventing |
The build-or-buy ladder#
you own → nothing gateway + data + serving + cluster + hardware
─────── ─────── ────── ───────── ───────── ──────────
hosted API ●
API + own gateway ●
+ own retrieval ●
managed GPU endpoints ●
own Kubernetes on cloud GPUs ●
own racks ●The triggers that justify each step down:
| Step | Do it when | Not before |
|---|---|---|
| Own a gateway | More than one application calls models | — do this first, always |
| Own the data layer | The product needs your documents or durable state | — almost always needed |
| Own serving | Open-weights or fine-tuned models are required; data must stay in your boundary; a steady load makes per-token pricing more expensive than GPU-hours | You have measured the steady load |
| Own the cluster | Several models and teams share GPUs; you need scheduling and quota control | One model on a managed endpoint would do |
| Own hardware | Utilisation is high and predictable for years; power and space are available | You have run rented GPUs at high occupancy for a year |
The break-even, roughly#
Self-hosting wins on unit cost only when the GPUs are busy. The comparison is one line:
API cost / day = tokens/day × price per token
self-hosted / day = GPUs × $/GPU-hour × 24 + people and platform overhead
A GPU that produces 2,500 output tokens/s is paid for whether it is busy or idle.
At 60% occupancy it yields ~130 M output tokens/day for ~$96/day → $0.74 per M
At 10% occupancy it yields ~22 M output tokens/day for ~$96/day → $4.44 per MOccupancy is the whole argument. Spiky, low-volume or experimental workloads belong on an API. Steady, high-volume workloads on a model that fits your hardware belong on your own GPUs. The overhead term is not small: a serving platform needs on-call engineers, and for a small fleet their cost can exceed the hardware’s. Cost per token does this arithmetic properly.
The order to build in#
Build from the top of the stack down, because each step works with rented layers beneath it:
- Gateway, pointing at hosted APIs. Immediately gives keys, quotas, attribution and a place to swap providers.
- Telemetry and evaluation. Usage events from the gateway; an evaluation suite. You now know your real token volumes — the numbers every later decision needs.
- Data layer. Ingestion, a vector index, session and task state.
- Serving, starting with one self-hosted model for one workload where the numbers justify it, added as another route behind the gateway.
- Cluster features: quotas, queues, autoscaling, as more teams arrive.
- Capacity commitments and hardware, once occupancy has been proven.
At every step the application code does not change, because it only ever talked to the gateway.
Control plane and data plane#
Separate the two from the start. The data plane carries requests: gateway, engines, retrieval. It must stay up without the control plane. The control plane decides what should be running: model registry, deployment controllers, autoscalers, policy. It can be down for minutes without users noticing — provided the data plane does not depend on it per request. Control plane versus data plane goes deeper.
Remember this#
- Five layers: compute, cluster, serving, data, platform.
- Decide which layer you start owning; own no more than a measured reason demands.
- Self-hosting is an occupancy bet, plus people.
- Build top-down: gateway, telemetry, data, serving, cluster, hardware.
- Keep the data plane independent of the control plane.
Try it#
- Place your current system on the ladder. What would have to be true to move one step down?
- Recompute the break-even for a GPU at $2.50/hour producing 1,200 tokens/s. At what occupancy does it match an API priced at $2 per million output tokens?
- List the control-plane components of a system you run. Which of them, if down, would stop user requests? Should it?
Check yourself#
- Why is the gateway the first thing to own?
- What single number decides whether self-hosting is cheaper than an API?
- Why can the control plane be less available than the data plane?