Pidoku

Deployment and Infrastructure

Intermediate 50 min Difficulty 3/5 Lesson 05 of 07

Prerequisites The Lifecycle Map

The idea in one minute#

Deployment security for AI is ordinary cloud and Kubernetes hardening applied to an unusual workload: one that holds extremely valuable assets (weights, prompts, customer context), runs on scarce shared hardware (GPUs), keeps state in places classic services do not (KV caches, vector indexes), and — when it is an agent — initiates outbound connections on the instructions of a model. Four things need designing: identity and secrets, network boundaries, isolation between tenants, and protection of the model and data while in use. Almost every item is a deterministic control, which makes this stage the most reliable place to bound damage.

A picture#

flowchart TB
  NET[":cloudflare: <b>Edge</b><br/><small>TLS, WAF, DDoS</small>"] --> GW[":envoyproxy: <b>AI gateway</b><br/><small>the only public endpoint</small>"]
  subgraph CLUSTER["Cluster: default-deny network policy"]
    direction LR
    subgraph NS1["Namespace: agents"]
      AG[":i-bot: <b>Agent runtime</b><br/><small>workload identity</small>"]
    end
    subgraph NS2["Namespace: serving"]
      SRV[":vllm: <b>Model servers</b><br/><small>no internet egress</small>"]
    end
    subgraph NS3["Namespace: sandboxes"]
      SB[":gvisor: <b>Sandboxes</b><br/><small>separate node pool</small>"]
    end
    subgraph NS4["Namespace: data"]
      DB[(":postgresql: <b>State, index</b><br/><small>encrypted, tenant-filtered</small>")]
    end
  end
  GW --> AG
  AG --> SRV
  AG --> SB
  AG --> DB
  AG --> VAULT[":vault: <b>Secrets and tokens</b><br/><small>short-lived, per task</small>"]
  SB --> EG[":cilium: <b>Egress proxy</b><br/><small>allowlist</small>"]
  SRV --> GPU[(":nvidia: <b>GPU nodes</b><br/><small>per-tenant pools where required</small>")]
  SPIRE[":spiffe: <b>Workload identity</b><br/><small>mTLS between services</small>"] -.-> AG
  SPIRE -.-> SRV
  class NET,GW queue
  class AG,SRV compute
  class SB,EG warn
  class DB,GPU,VAULT memory
  class SPIRE io

How it really works#

Identity and secrets#

No long-lived keys. The most common AI security incident is still a leaked credential — a model API key in a repository, a notebook, a client-side bundle or an agent’s configuration file.

RuleHow
Provider keys exist only in the gatewayApplications receive gateway keys scoped to a project and budget
Workloads authenticate as themselvesWorkload identity — SPIFFE/SPIRE, cloud workload identity — instead of shared secrets; mutual TLS between services
Secrets are fetched at run timeFrom a secret manager, never baked into images, prompts or environment files in a repository
Credentials are short-livedMinutes to hours; rotation is automatic
Nothing secret enters a model contextTools add credentials in code, after the model has proposed the call
Leaks are assumedSecret scanning in repositories and logs; provider-side spend limits; a practised revocation procedure

The fifth rule is specific to AI and absolute: a credential in the context can be printed, sent or logged by a manipulated model.

Network boundaries#

Start from deny-all and open what is needed.

ComponentInboundOutbound
AI gatewayThe internet, through the edgeModel providers; internal model servers
Agent runtimeThe gateway onlyGateway (for model calls), tool gateway, state stores
Model serversGateway and router onlyNone. Weights come from the internal registry
SandboxesTheir controller onlyAn allowlist, through a proxy; never the cluster network or cloud metadata
MCP serversThe tool gateway onlyThe one system each fronts
Data storesNamed services onlyNone
Training jobsNoneDataset and artifact stores

Three specifics: block the cloud metadata endpoint from anything that runs model-chosen code or requests, since it hands out cloud credentials; keep admin and debug interfaces of inference servers and vector databases off public networks — exposed, unauthenticated ones are found by scanners constantly; and put inference engines’ own HTTP ports behind the gateway, because most engines ship without authentication.

Isolation between tenants#

A shared AI platform has more places for one tenant’s data to meet another’s than a classic service.

Shared resourceLeak pathIsolation
Vector indexA missing filter returns another tenant’s chunksMandatory tenant filter enforced by the service layer and tested; separate collections for strict tenants
Prompt / KV prefix cacheTiming reveals that another tenant sent the same prefixPer-tenant cache partitioning or salting
Semantic or response cacheOne tenant receives another’s answerNever shared across tenants
Batching in the engineSequences from different tenants share a batch — isolated in principle, exposed by engine bugsDedicated replicas for tenants that require it
GPU memoryResidual data between workloads; side channels on shared devicesWhole-GPU allocation per tenant, or hardware partitioning (MIG) rather than time-slicing, for sensitive workloads
SandboxesKernel escape; shared file systemsOne sandbox per task; gVisor or microVM; separate node pool
Logs and tracesEngineers reading another tenant’s promptsContent off by default; per-tenant access control
Fine-tuned adaptersOne tenant’s adapter served to anotherAdapter access bound to tenant identity at the router

Decide tenancy tiers explicitly: shared (logical isolation, filters and partitioned caches), dedicated (own replicas and node pools) and confidential (hardware-enforced). See security and isolation for the serving-side detail.

Protecting models and data in use#

Encryption at rest and in transit is the baseline. Data in use — prompts and weights in memory during inference — is exposed to whoever controls the host: the cloud operator, a compromised hypervisor, an administrator.

Confidential computing closes that gap with hardware:

  • A trusted execution environment — a confidential VM on AMD SEV-SNP or Intel TDX — encrypts memory so the host cannot read it.
  • Confidential GPUs extend this to the accelerator: the GPU runs in a protected mode and the CPU-to-GPU channel is encrypted. NVIDIA supports this on Hopper and Blackwell generations.
  • Attestation is what makes it useful. The hardware signs a measurement of exactly what was loaded — firmware, kernel, container image, GPU state — and a key-release service hands over the model decryption key or the tenant’s data key only if the measurement matches an approved value.
  • Confidential Containers bring the pattern to Kubernetes pods.

Measured overhead on inference is modest — single-digit percent in published benchmarks — with more operational complexity than performance cost. Use it where the threat model includes the infrastructure operator: regulated data, a customer who will not trust your cloud, or weights valuable enough to protect from your own administrators.

Hardening the workloads#

The Kubernetes basics apply and matter more than usual:

  • Pod security: non-root, read-only root file system, dropped capabilities, no privileged containers — with the known exception that GPU device access needs careful, minimal grants.
  • Admission policy: only signed images and signed models from the internal registry; no latest tags; resource limits required.
  • Separate node pools for sandboxes, serving, data and system workloads, so a sandbox escape lands on a node with nothing valuable.
  • Patch the stack: GPU drivers, container toolkits and inference engines have had serious vulnerabilities, including container escapes through GPU tooling. They are internet-era software running with high privilege.
  • Runtime detection: eBPF-based tools (Falco, Tetragon) alert on unexpected processes, file access or network connections in serving and agent pods — a model server that starts a shell is an incident.
  • Infrastructure as code, reviewed, with drift detection.

Protecting the weights#

For proprietary or fine-tuned models, weights are the crown jewels and are simply large files.

  • Encrypt at rest, with keys in a KMS; restrict who and what can read the registry.
  • Serving pods mount weights read-only and have no route out of the cluster.
  • Alert on unusually large reads or transfers from model storage.
  • Limit what the API exposes — full log-probabilities and unlimited query volume make distillation easier.
  • For the highest sensitivity: attestation-gated key release, so weights decrypt only inside an approved environment.

A deployment checklist#

  • One public endpoint: the gateway. Engines, vector stores and admin UIs are private.
  • No provider keys outside the gateway; no secrets in prompts, images or repositories.
  • Workload identity and mutual TLS between services.
  • Default-deny network policy; model servers and data stores have no internet egress.
  • Sandboxes on their own nodes with allowlisted egress and no metadata access.
  • Tenant isolation defined per shared resource, including caches.
  • Only signed images and models from the internal registry, by digest.
  • Drivers, engines and toolkits on a patch schedule.
  • Runtime detection on serving, agent and sandbox pods.
  • Weights encrypted, access-logged, with transfer alerts.

Remember this#

  • The leaked key is still the commonest incident: short-lived, scoped, never in a context.
  • Default-deny networking; model servers need no internet.
  • Caches and indexes are tenant-isolation boundaries, not just performance features.
  • Confidential computing protects data in use and depends on attestation.
  • Inference engines and GPU tooling are privileged software to be kept private and patched.

Try it#

  1. Draw the network table for a system you run. Which component has egress it does not need?
  2. List every cache in your stack and state its tenant scope.
  3. Decide which, if any, of your tenants or datasets justify dedicated or confidential tiers, and why.

Check yourself#

  1. Why must credentials never appear in a model’s context?
  2. How can a shared prefix cache leak information between tenants?
  3. What does attestation add to memory encryption?

Sources#

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom