The idea in one minute#
Deployment security for AI is ordinary cloud and Kubernetes hardening applied to an unusual workload: one that holds extremely valuable assets (weights, prompts, customer context), runs on scarce shared hardware (GPUs), keeps state in places classic services do not (KV caches, vector indexes), and — when it is an agent — initiates outbound connections on the instructions of a model. Four things need designing: identity and secrets, network boundaries, isolation between tenants, and protection of the model and data while in use. Almost every item is a deterministic control, which makes this stage the most reliable place to bound damage.
A picture#
flowchart TB
NET[":cloudflare: <b>Edge</b><br/><small>TLS, WAF, DDoS</small>"] --> GW[":envoyproxy: <b>AI gateway</b><br/><small>the only public endpoint</small>"]
subgraph CLUSTER["Cluster: default-deny network policy"]
direction LR
subgraph NS1["Namespace: agents"]
AG[":i-bot: <b>Agent runtime</b><br/><small>workload identity</small>"]
end
subgraph NS2["Namespace: serving"]
SRV[":vllm: <b>Model servers</b><br/><small>no internet egress</small>"]
end
subgraph NS3["Namespace: sandboxes"]
SB[":gvisor: <b>Sandboxes</b><br/><small>separate node pool</small>"]
end
subgraph NS4["Namespace: data"]
DB[(":postgresql: <b>State, index</b><br/><small>encrypted, tenant-filtered</small>")]
end
end
GW --> AG
AG --> SRV
AG --> SB
AG --> DB
AG --> VAULT[":vault: <b>Secrets and tokens</b><br/><small>short-lived, per task</small>"]
SB --> EG[":cilium: <b>Egress proxy</b><br/><small>allowlist</small>"]
SRV --> GPU[(":nvidia: <b>GPU nodes</b><br/><small>per-tenant pools where required</small>")]
SPIRE[":spiffe: <b>Workload identity</b><br/><small>mTLS between services</small>"] -.-> AG
SPIRE -.-> SRV
class NET,GW queue
class AG,SRV compute
class SB,EG warn
class DB,GPU,VAULT memory
class SPIRE ioHow it really works#
Identity and secrets#
No long-lived keys. The most common AI security incident is still a leaked credential — a model API key in a repository, a notebook, a client-side bundle or an agent’s configuration file.
| Rule | How |
|---|---|
| Provider keys exist only in the gateway | Applications receive gateway keys scoped to a project and budget |
| Workloads authenticate as themselves | Workload identity — SPIFFE/SPIRE, cloud workload identity — instead of shared secrets; mutual TLS between services |
| Secrets are fetched at run time | From a secret manager, never baked into images, prompts or environment files in a repository |
| Credentials are short-lived | Minutes to hours; rotation is automatic |
| Nothing secret enters a model context | Tools add credentials in code, after the model has proposed the call |
| Leaks are assumed | Secret scanning in repositories and logs; provider-side spend limits; a practised revocation procedure |
The fifth rule is specific to AI and absolute: a credential in the context can be printed, sent or logged by a manipulated model.
Network boundaries#
Start from deny-all and open what is needed.
| Component | Inbound | Outbound |
|---|---|---|
| AI gateway | The internet, through the edge | Model providers; internal model servers |
| Agent runtime | The gateway only | Gateway (for model calls), tool gateway, state stores |
| Model servers | Gateway and router only | None. Weights come from the internal registry |
| Sandboxes | Their controller only | An allowlist, through a proxy; never the cluster network or cloud metadata |
| MCP servers | The tool gateway only | The one system each fronts |
| Data stores | Named services only | None |
| Training jobs | None | Dataset and artifact stores |
Three specifics: block the cloud metadata endpoint from anything that runs model-chosen code or requests, since it hands out cloud credentials; keep admin and debug interfaces of inference servers and vector databases off public networks — exposed, unauthenticated ones are found by scanners constantly; and put inference engines’ own HTTP ports behind the gateway, because most engines ship without authentication.
Isolation between tenants#
A shared AI platform has more places for one tenant’s data to meet another’s than a classic service.
| Shared resource | Leak path | Isolation |
|---|---|---|
| Vector index | A missing filter returns another tenant’s chunks | Mandatory tenant filter enforced by the service layer and tested; separate collections for strict tenants |
| Prompt / KV prefix cache | Timing reveals that another tenant sent the same prefix | Per-tenant cache partitioning or salting |
| Semantic or response cache | One tenant receives another’s answer | Never shared across tenants |
| Batching in the engine | Sequences from different tenants share a batch — isolated in principle, exposed by engine bugs | Dedicated replicas for tenants that require it |
| GPU memory | Residual data between workloads; side channels on shared devices | Whole-GPU allocation per tenant, or hardware partitioning (MIG) rather than time-slicing, for sensitive workloads |
| Sandboxes | Kernel escape; shared file systems | One sandbox per task; gVisor or microVM; separate node pool |
| Logs and traces | Engineers reading another tenant’s prompts | Content off by default; per-tenant access control |
| Fine-tuned adapters | One tenant’s adapter served to another | Adapter access bound to tenant identity at the router |
Decide tenancy tiers explicitly: shared (logical isolation, filters and partitioned caches), dedicated (own replicas and node pools) and confidential (hardware-enforced). See security and isolation for the serving-side detail.
Protecting models and data in use#
Encryption at rest and in transit is the baseline. Data in use — prompts and weights in memory during inference — is exposed to whoever controls the host: the cloud operator, a compromised hypervisor, an administrator.
Confidential computing closes that gap with hardware:
- A trusted execution environment — a confidential VM on AMD SEV-SNP or Intel TDX — encrypts memory so the host cannot read it.
- Confidential GPUs extend this to the accelerator: the GPU runs in a protected mode and the CPU-to-GPU channel is encrypted. NVIDIA supports this on Hopper and Blackwell generations.
- Attestation is what makes it useful. The hardware signs a measurement of exactly what was loaded — firmware, kernel, container image, GPU state — and a key-release service hands over the model decryption key or the tenant’s data key only if the measurement matches an approved value.
- Confidential Containers bring the pattern to Kubernetes pods.
Measured overhead on inference is modest — single-digit percent in published benchmarks — with more operational complexity than performance cost. Use it where the threat model includes the infrastructure operator: regulated data, a customer who will not trust your cloud, or weights valuable enough to protect from your own administrators.
Hardening the workloads#
The Kubernetes basics apply and matter more than usual:
- Pod security: non-root, read-only root file system, dropped capabilities, no privileged containers — with the known exception that GPU device access needs careful, minimal grants.
- Admission policy: only signed images and signed models from the internal registry; no
latesttags; resource limits required. - Separate node pools for sandboxes, serving, data and system workloads, so a sandbox escape lands on a node with nothing valuable.
- Patch the stack: GPU drivers, container toolkits and inference engines have had serious vulnerabilities, including container escapes through GPU tooling. They are internet-era software running with high privilege.
- Runtime detection: eBPF-based tools (Falco, Tetragon) alert on unexpected processes, file access or network connections in serving and agent pods — a model server that starts a shell is an incident.
- Infrastructure as code, reviewed, with drift detection.
Protecting the weights#
For proprietary or fine-tuned models, weights are the crown jewels and are simply large files.
- Encrypt at rest, with keys in a KMS; restrict who and what can read the registry.
- Serving pods mount weights read-only and have no route out of the cluster.
- Alert on unusually large reads or transfers from model storage.
- Limit what the API exposes — full log-probabilities and unlimited query volume make distillation easier.
- For the highest sensitivity: attestation-gated key release, so weights decrypt only inside an approved environment.
A deployment checklist#
- One public endpoint: the gateway. Engines, vector stores and admin UIs are private.
- No provider keys outside the gateway; no secrets in prompts, images or repositories.
- Workload identity and mutual TLS between services.
- Default-deny network policy; model servers and data stores have no internet egress.
- Sandboxes on their own nodes with allowlisted egress and no metadata access.
- Tenant isolation defined per shared resource, including caches.
- Only signed images and models from the internal registry, by digest.
- Drivers, engines and toolkits on a patch schedule.
- Runtime detection on serving, agent and sandbox pods.
- Weights encrypted, access-logged, with transfer alerts.
Remember this#
- The leaked key is still the commonest incident: short-lived, scoped, never in a context.
- Default-deny networking; model servers need no internet.
- Caches and indexes are tenant-isolation boundaries, not just performance features.
- Confidential computing protects data in use and depends on attestation.
- Inference engines and GPU tooling are privileged software to be kept private and patched.
Try it#
- Draw the network table for a system you run. Which component has egress it does not need?
- List every cache in your stack and state its tenant scope.
- Decide which, if any, of your tenants or datasets justify dedicated or confidential tiers, and why.
Check yourself#
- Why must credentials never appear in a model’s context?
- How can a shared prefix cache leak information between tenants?
- What does attestation add to memory encryption?