Pidoku

The 2026 Landscape

Foundations 45 min Difficulty 2/5 Lesson 03 of 03

Prerequisites What a Model Changes

The idea in one minute#

An AI system in 2026 is seven layers, and each layer has settled on a small number of real options. From the bottom: compute, cluster, serving, gateway, data and context, orchestration, and across all of them observability and security. Two open protocols connect the upper layers — MCP between an agent and its tools, A2A between agents — and both are now governed by one neutral foundation. If you can place a tool on this map, you know what it replaces and what it does not.

Everything dated in this lesson was checked on 4 October 2026.

A picture#

flowchart TB
  subgraph APP["Application and orchestration"]
    direction TB
    L1[":langgraph: <b>LangGraph</b><br/><small>graph runtime, checkpoints</small>"]
    L2[":claude: <b>Claude Agent SDK</b><br/><small>agent with a computer</small>"]
    L3[":openai: <b>OpenAI Agents SDK</b><br/><small>handoffs, sandboxes</small>"]
    L4[":temporal: <b>Temporal</b><br/><small>durable execution</small>"]
  end
  subgraph PROTO["Protocols"]
    direction TB
    P1[":modelcontextprotocol: <b>MCP</b><br/><small>agent to tools</small>"]
    P2[":a2a: <b>A2A</b><br/><small>agent to agent</small>"]
  end
  subgraph CTX["Data and context"]
    direction TB
    D1[":postgresql: <b>pgvector</b>"]
    D2[":qdrant: <b>Qdrant</b>"]
    D3[":milvus: <b>Milvus</b>"]
    D4[":redis: <b>Redis</b><br/><small>cache, sessions</small>"]
  end
  subgraph GATE["Gateway"]
    direction TB
    G1[":envoyproxy: <b>Envoy AI Gateway</b>"]
    G2[":agentgateway: <b>agentgateway</b>"]
    G3[":litellm: <b>LiteLLM</b>"]
  end
  subgraph SERVE["Serving"]
    direction TB
    S1[":vllm: <b>vLLM</b>"]
    S2[":sglang: <b>SGLang</b>"]
    S3[":nvidia: <b>Dynamo</b><br/><small>distributed serving</small>"]
    S4[":llm-d: <b>llm-d</b><br/><small>Kubernetes-native</small>"]
  end
  subgraph CLUSTER["Cluster"]
    direction TB
    K1[":kubernetes: <b>Kubernetes</b><br/><small>DRA, Gateway API</small>"]
    K2[":kueue: <b>Kueue</b><br/><small>quota, queues</small>"]
    K3[":ray: <b>Ray</b>"]
  end
  subgraph HW["Compute"]
    direction TB
    H1[":nvidia: <b>NVIDIA GPUs</b>"]
    H2[":amd: <b>AMD Instinct</b>"]
    H3[":googlecloud: <b>Cloud accelerators</b>"]
  end
  APP --> PROTO --> CTX
  APP --> GATE --> SERVE --> CLUSTER --> HW
  class L1,L2,L3,L4 compute
  class P1,P2 io
  class D1,D2,D3,D4 memory
  class G1,G2,G3 queue
  class S1,S2,S3,S4 compute
  class K1,K2,K3 neutral
  class H1,H2,H3 memory

Hosted model APIs enter at the gateway layer: from the application’s point of view, a frontier model behind an API and a model on your own GPUs are two routes behind the same gateway.

How it really works#

The seven layers#

LayerIts jobCommon choices in 2026Deep dive
ComputeRun the arithmeticNVIDIA Blackwell and Rubin, AMD Instinct, cloud-designed acceleratorsGPU generations
ClusterHand out accelerators, queue work, share fairlyKubernetes with DRA, Kueue, the NVIDIA GPU Operator; Slurm for training; Ray for Python-native jobsThe Cluster
ServingTurn weights into tokens efficientlyvLLM, SGLang, TensorRT-LLM; llm-d or NVIDIA Dynamo to coordinate many replicasThe Serving Layer
GatewayOne entry point: auth, token limits, routing, fallbackEnvoy AI Gateway, agentgateway, LiteLLM, Kong, cloud gatewaysModel Access and Gateways
Data and contextStore and find what the model should seePostgres with pgvector, Qdrant, Milvus, OpenSearch; Redis for sessions and caches; object storage for documents and weightsContext and Retrieval
OrchestrationRun the loop, keep state, survive crashesLangGraph, Claude Agent SDK, OpenAI Agents SDK, Google ADK, Microsoft Agent Framework; Temporal or Restate underneath for durabilityAgentic Systems
Observability and securitySee it and constrain itOpenTelemetry, Prometheus, Grafana, Langfuse; sandboxes, policy engines, guardrail modelsObservability, Security

The standards that connect the layers#

Standards matter to a designer because they decide where you can swap a component.

StandardConnectsState on 4 October 2026
OpenAI-compatible HTTP APIApplication ↔ any model serverDe facto. Every engine and gateway speaks it, which is what makes “hosted or self-hosted” a routing decision
MCP (Model Context Protocol)Agent ↔ tools and dataSpecification 2026-07-28: the protocol core is now stateless — no initialize handshake, no session header — so servers sit behind an ordinary load balancer. Method and tool names travel in HTTP headers so gateways can route and meter without parsing bodies. Tasks, MCP Apps and Enterprise-Managed Authorization are official extensions
A2A (Agent2Agent)Agent ↔ agentv1.0 since March 2026. Agents publish a signed Agent Card describing what they can do; tasks are long-running and resumable
Gateway API Inference ExtensionGateway ↔ model servers on KubernetesGenerally available. InferencePool groups model-server pods; an endpoint picker routes on queue depth, KV-cache use and loaded adapters instead of round-robin
Dynamic Resource AllocationPod ↔ acceleratorCore API stable since Kubernetes 1.34; 1.36 added more of the device features. Requests describe a device by attribute rather than by count
OpenTelemetry GenAI conventionsEverything ↔ telemetry backendStill marked Development, now in their own repository; widely emitted anyway. Pin the version you depend on

MCP and A2A are both projects of the Agentic AI Foundation, formed under the Linux Foundation in December 2025 with MCP, goose and AGENTS.md as founding projects; A2A joined in August 2026. For a designer the consequence is that neither protocol is a single vendor’s to change.

Three shifts that shaped this map#

From model-centric to system-centric. In 2023 the question was “which model?”. Models are now good enough that most quality differences between products come from what surrounds the model: the context it is given, the tools it has, the evaluation that gates a release.

From stateful to stateless at the edges. MCP dropping its session, inference routing moving into the gateway, agents checkpointing to external stores — each makes a component replaceable and horizontally scalable. State concentrates in a few stores you choose on purpose.

From chat to work. Agents that run for minutes or hours turned serving into a background workload: far more input tokens than output tokens, long idle periods waiting on tools, and sandboxes as a new unit of compute. Kubernetes now has an Agent Sandbox API for exactly that unit — a stateful, single-tenant environment that can idle and resume.

How to place a new tool#

When something new is announced, ask four questions and you will know what it is:

  1. Which layer? If it claims three, find the one it is actually good at.
  2. What does it replace, and what does it sit beside? A gateway replaces hand-written retry code; it does not replace an inference engine.
  3. Which standard does it speak on each side? Standards in and out mean you can remove it later.
  4. Where does it keep state? Anything holding state is a migration; anything stateless is a configuration change.

Remember this#

  • Seven layers: compute, cluster, serving, gateway, data and context, orchestration, and observability plus security across all.
  • Hosted and self-hosted models are two routes behind one gateway.
  • MCP connects agents to tools, A2A connects agents to agents; both sit under one foundation.
  • Prefer components that speak a standard on both sides and hold no state.

Try it#

  1. Take the stack of a product you know and write one tool per layer. Which layers are provided by a vendor on your behalf?
  2. Pick any recently announced AI infrastructure project and answer the four placement questions for it.

Check yourself#

  1. Why does a stateless MCP core matter to someone running servers behind a load balancer?
  2. What does an inference-aware gateway look at that a round-robin load balancer does not?
  3. Which layer would you change to move a workload from a hosted API to your own GPUs, and which layers should not notice?

Sources#

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom