The idea in one minute#
An AI system in 2026 is seven layers, and each layer has settled on a small number of real options. From the bottom: compute, cluster, serving, gateway, data and context, orchestration, and across all of them observability and security. Two open protocols connect the upper layers — MCP between an agent and its tools, A2A between agents — and both are now governed by one neutral foundation. If you can place a tool on this map, you know what it replaces and what it does not.
Everything dated in this lesson was checked on 4 October 2026.
A picture#
flowchart TB
subgraph APP["Application and orchestration"]
direction TB
L1[":langgraph: <b>LangGraph</b><br/><small>graph runtime, checkpoints</small>"]
L2[":claude: <b>Claude Agent SDK</b><br/><small>agent with a computer</small>"]
L3[":openai: <b>OpenAI Agents SDK</b><br/><small>handoffs, sandboxes</small>"]
L4[":temporal: <b>Temporal</b><br/><small>durable execution</small>"]
end
subgraph PROTO["Protocols"]
direction TB
P1[":modelcontextprotocol: <b>MCP</b><br/><small>agent to tools</small>"]
P2[":a2a: <b>A2A</b><br/><small>agent to agent</small>"]
end
subgraph CTX["Data and context"]
direction TB
D1[":postgresql: <b>pgvector</b>"]
D2[":qdrant: <b>Qdrant</b>"]
D3[":milvus: <b>Milvus</b>"]
D4[":redis: <b>Redis</b><br/><small>cache, sessions</small>"]
end
subgraph GATE["Gateway"]
direction TB
G1[":envoyproxy: <b>Envoy AI Gateway</b>"]
G2[":agentgateway: <b>agentgateway</b>"]
G3[":litellm: <b>LiteLLM</b>"]
end
subgraph SERVE["Serving"]
direction TB
S1[":vllm: <b>vLLM</b>"]
S2[":sglang: <b>SGLang</b>"]
S3[":nvidia: <b>Dynamo</b><br/><small>distributed serving</small>"]
S4[":llm-d: <b>llm-d</b><br/><small>Kubernetes-native</small>"]
end
subgraph CLUSTER["Cluster"]
direction TB
K1[":kubernetes: <b>Kubernetes</b><br/><small>DRA, Gateway API</small>"]
K2[":kueue: <b>Kueue</b><br/><small>quota, queues</small>"]
K3[":ray: <b>Ray</b>"]
end
subgraph HW["Compute"]
direction TB
H1[":nvidia: <b>NVIDIA GPUs</b>"]
H2[":amd: <b>AMD Instinct</b>"]
H3[":googlecloud: <b>Cloud accelerators</b>"]
end
APP --> PROTO --> CTX
APP --> GATE --> SERVE --> CLUSTER --> HW
class L1,L2,L3,L4 compute
class P1,P2 io
class D1,D2,D3,D4 memory
class G1,G2,G3 queue
class S1,S2,S3,S4 compute
class K1,K2,K3 neutral
class H1,H2,H3 memoryHosted model APIs enter at the gateway layer: from the application’s point of view, a frontier model behind an API and a model on your own GPUs are two routes behind the same gateway.
How it really works#
The seven layers#
| Layer | Its job | Common choices in 2026 | Deep dive |
|---|---|---|---|
| Compute | Run the arithmetic | NVIDIA Blackwell and Rubin, AMD Instinct, cloud-designed accelerators | GPU generations |
| Cluster | Hand out accelerators, queue work, share fairly | Kubernetes with DRA, Kueue, the NVIDIA GPU Operator; Slurm for training; Ray for Python-native jobs | The Cluster |
| Serving | Turn weights into tokens efficiently | vLLM, SGLang, TensorRT-LLM; llm-d or NVIDIA Dynamo to coordinate many replicas | The Serving Layer |
| Gateway | One entry point: auth, token limits, routing, fallback | Envoy AI Gateway, agentgateway, LiteLLM, Kong, cloud gateways | Model Access and Gateways |
| Data and context | Store and find what the model should see | Postgres with pgvector, Qdrant, Milvus, OpenSearch; Redis for sessions and caches; object storage for documents and weights | Context and Retrieval |
| Orchestration | Run the loop, keep state, survive crashes | LangGraph, Claude Agent SDK, OpenAI Agents SDK, Google ADK, Microsoft Agent Framework; Temporal or Restate underneath for durability | Agentic Systems |
| Observability and security | See it and constrain it | OpenTelemetry, Prometheus, Grafana, Langfuse; sandboxes, policy engines, guardrail models | Observability, Security |
The standards that connect the layers#
Standards matter to a designer because they decide where you can swap a component.
| Standard | Connects | State on 4 October 2026 |
|---|---|---|
| OpenAI-compatible HTTP API | Application ↔ any model server | De facto. Every engine and gateway speaks it, which is what makes “hosted or self-hosted” a routing decision |
| MCP (Model Context Protocol) | Agent ↔ tools and data | Specification 2026-07-28: the protocol core is now stateless — no initialize handshake, no session header — so servers sit behind an ordinary load balancer. Method and tool names travel in HTTP headers so gateways can route and meter without parsing bodies. Tasks, MCP Apps and Enterprise-Managed Authorization are official extensions |
| A2A (Agent2Agent) | Agent ↔ agent | v1.0 since March 2026. Agents publish a signed Agent Card describing what they can do; tasks are long-running and resumable |
| Gateway API Inference Extension | Gateway ↔ model servers on Kubernetes | Generally available. InferencePool groups model-server pods; an endpoint picker routes on queue depth, KV-cache use and loaded adapters instead of round-robin |
| Dynamic Resource Allocation | Pod ↔ accelerator | Core API stable since Kubernetes 1.34; 1.36 added more of the device features. Requests describe a device by attribute rather than by count |
| OpenTelemetry GenAI conventions | Everything ↔ telemetry backend | Still marked Development, now in their own repository; widely emitted anyway. Pin the version you depend on |
MCP and A2A are both projects of the Agentic AI Foundation, formed under the Linux Foundation in December 2025 with MCP, goose and AGENTS.md as founding projects; A2A joined in August 2026. For a designer the consequence is that neither protocol is a single vendor’s to change.
Three shifts that shaped this map#
From model-centric to system-centric. In 2023 the question was “which model?”. Models are now good enough that most quality differences between products come from what surrounds the model: the context it is given, the tools it has, the evaluation that gates a release.
From stateful to stateless at the edges. MCP dropping its session, inference routing moving into the gateway, agents checkpointing to external stores — each makes a component replaceable and horizontally scalable. State concentrates in a few stores you choose on purpose.
From chat to work. Agents that run for minutes or hours turned serving into a background workload: far more input tokens than output tokens, long idle periods waiting on tools, and sandboxes as a new unit of compute. Kubernetes now has an Agent Sandbox API for exactly that unit — a stateful, single-tenant environment that can idle and resume.
How to place a new tool#
When something new is announced, ask four questions and you will know what it is:
- Which layer? If it claims three, find the one it is actually good at.
- What does it replace, and what does it sit beside? A gateway replaces hand-written retry code; it does not replace an inference engine.
- Which standard does it speak on each side? Standards in and out mean you can remove it later.
- Where does it keep state? Anything holding state is a migration; anything stateless is a configuration change.
Remember this#
- Seven layers: compute, cluster, serving, gateway, data and context, orchestration, and observability plus security across all.
- Hosted and self-hosted models are two routes behind one gateway.
- MCP connects agents to tools, A2A connects agents to agents; both sit under one foundation.
- Prefer components that speak a standard on both sides and hold no state.
Try it#
- Take the stack of a product you know and write one tool per layer. Which layers are provided by a vendor on your behalf?
- Pick any recently announced AI infrastructure project and answer the four placement questions for it.
Check yourself#
- Why does a stateless MCP core matter to someone running servers behind a load balancer?
- What does an inference-aware gateway look at that a round-robin load balancer does not?
- Which layer would you change to move a workload from a hosted API to your own GPUs, and which layers should not notice?