The idea in one minute#
Tools are how an agent touches the world, and the Model Context Protocol made tools easy to publish and easy to install. That convenience moved the supply-chain problem into the agent: an MCP server is third-party code and third-party text, both arriving in a privileged place. Securing tools has three sides. As a consumer, treat each server as a supplier: vet, pin, isolate and monitor it. As a publisher, build servers that authenticate callers, validate input and expose the least. As a platform, put a gateway between agents and servers so that approval, authorization and audit happen in one place.
Incident references in this lesson were checked on 4 October 2026.
A picture#
flowchart LR
AG[":i-bot: <b>Agents</b>"] --> TG
subgraph TG["Tool gateway"]
direction TB
T1[":i-list-checks: <b>Catalogue</b><br/><small>approved servers, pinned versions,<br/>description hashes</small>"]
T2[":i-fingerprint: <b>AuthN / AuthZ</b><br/><small>per user, per tool,<br/>audience-bound tokens</small>"]
T3[":i-shield-check: <b>Inspect</b><br/><small>arguments in, results out</small>"]
T4[":i-scroll-text: <b>Audit and rate limit</b>"]
T1 --> T2 --> T3 --> T4
end
TG --> S1[":github: <b>Source control MCP</b><br/><small>first party</small>"]
TG --> S2[":jira: <b>Tracker MCP</b><br/><small>vendor, remote</small>"]
TG --> S3[":i-box: <b>Community MCP</b><br/><small>sandboxed, restricted egress</small>"]
REG[":modelcontextprotocol: <b>Public registries</b>"] -.->|"review, scan, pin"| T1
S3 -.->|"description changed"| ALERT[":i-siren: <b>Re-review</b>"]
class AG compute
class T1,T2,T3,T4 queue
class S1,S2 io
class S3 warn
class REG neutral
class ALERT warnHow it really works#
What has actually gone wrong#
The incident record since 2025 shows the same few failures repeatedly:
| Failure | What happened |
|---|---|
| Malicious package | A legitimate-looking email MCP package shipped an update that silently copied every message to an outside address — the first tracked malicious MCP server, September 2025 |
| Tool poisoning | Demonstrations of a harmless-looking tool whose description instructed the model to read the user’s SSH key and pass it along as a parameter |
| Unauthenticated servers | Internet scans finding hundreds of MCP servers reachable with no authentication at all |
| Classic bugs in servers | Dozens of CVEs filed within weeks in early 2026 — command injection the largest group — plus an authentication bypass in a widely used administration tool’s MCP endpoint that was exploited in the wild |
| Auto-executed configuration | Flaws in IDE extensions where merely opening a repository containing a crafted MCP configuration file ran attacker commands; and a 2026 campaign that planted such files across public repositories to harvest developer credentials |
| SDK-level weaknesses | Vulnerabilities reported in protocol implementations themselves, affecting many downstream servers |
| Cross-tool data flows | An agent with a “read private repository” tool and a “comment on public issue” tool being steered, by text in an issue, to copy private code into the public comment |
Two lessons. First, most of this is ordinary software insecurity — missing authentication, unvalidated input — in code written quickly. Second, the uniquely new part is that text from a tool is input to a model that holds other tools.
As a consumer: vetting a server#
Apply the six supply-chain questions.
origin Who publishes it? A vendor for their own product, or an unknown third party?
integrity Is the artifact signed or at least pinned by digest? Is the source public?
content Read the tool descriptions in full — they are prompts.
Read the code, or at minimum what it executes, reads and connects to.
Run it in a sandbox and observe file and network activity.
pinning Exact version; hash of the tool list and descriptions recorded at approval.
privilege What credentials does it receive? What can it reach on the network and disk?
removal Can we disable it for every agent at once?Prefer, in order: first-party servers you build; the system vendor’s own official server; well-maintained open-source servers you have reviewed; and last, anything else — sandboxed.
Pin descriptions, not just versions. A remote server can change what its tools say at any time. Record a hash of the tool list at approval and compare on every connection; a change suspends the server pending review. This is the defence against rug pulls. The 2026-07-28 specification’s cacheable tool lists make the check cheap.
As a consumer: containing a server#
- Local (
stdio) servers run with the launching user’s privileges. Run them in a sandbox with a restricted file system and egress allowlist, and never auto-start one from a repository’s configuration. - Remote servers receive only an audience-bound, scoped, short-lived token for that server — never a general credential, and never another server’s token.
- Limit the tool set per task. Do not connect every server to every agent. Each connected server adds descriptions to the context and tools to the attack surface.
- Namespace tools by server, so one cannot shadow another’s name, and so descriptions from one server cannot plausibly instruct the use of another’s.
- Treat results as untrusted. They pass through the input rail, set the session taint, and are size-limited.
- Watch for dangerous pairs. A server that reads private data and a server that can publish, connected to the same agent, complete the trifecta. Evaluate combinations, not only servers.
As a platform: the tool gateway#
Putting a gateway between agents and MCP servers centralises what would otherwise be re-implemented per agent.
| Function | Detail |
|---|---|
| Catalogue | The only servers agents can reach; each with an owner, a risk tier, a pinned version and description hash |
| Authentication | Verifies the agent’s identity and the user’s delegated token |
| Authorization | Per user, per agent, per tool; since 2026-07-28 the method and tool name are in HTTP headers, so the gateway decides without parsing bodies |
| Token brokering | Exchanges the incoming token for one bound to the target server; no passthrough |
| Inspection | Argument validation; secret and PII detection outbound; injection screening and size limits on results |
| Rate limits and budgets | Per tool, per agent, per tenant |
| Audit | One record per call, joined to the task’s trace |
| Kill switch | Disable a server or a tool for everyone, immediately |
Gateways that speak MCP natively — agentgateway, Envoy AI Gateway and others — provide most of this, and with a stateless protocol core they scale like any HTTP proxy.
As a publisher: building a safe server#
If you write MCP servers, you are writing an API that will be called with arguments chosen by a manipulated model.
- Authenticate every request. OAuth 2.1 as a resource server; validate issuer, audience and expiry; reject tokens minted for another audience. No unauthenticated remote servers, and no listening on all interfaces by default.
- Authorise as the user. Enforce the end user’s permissions in the backing system; do not give the server one all-powerful account.
- Validate all input. Arguments are attacker-influenced. No string concatenation into shells, SQL, file paths or URLs — the command-injection CVEs are exactly this. Use parameterised calls, allowlists, and path confinement.
- Expose narrow tools.
get_invoice(id)rather thanrun_query(sql). Separate read tools from write tools so clients can grant them separately. Mark destructive tools as such in their annotations. - Write honest, minimal descriptions. Describe the tool; do not instruct the model about anything else. Clients may flag descriptions that mention other tools or files.
- Bound outputs. Paginate; cap sizes; return structured data. Label fields that contain user-generated text so clients can treat them as untrusted.
- Do not pass tokens through. Obtain your own credential for upstream APIs.
- Make writes idempotent and support dry-run where the effect is significant.
- Log caller, user, tool and outcome.
- Ship like a product: pinned dependencies, signed releases, an SBOM, a security contact and prompt patches.
Skills, plugins and instruction files#
Everything above applies to non-MCP extensions that carry instructions: skills, plugins, rules files, prompt packs, agent configuration checked into repositories.
- Installing one lets its author write to your agent’s instructions. Read it first.
- Keep an approved internal catalogue; pin by commit or digest.
- Bundled scripts are code: review, and run in the sandbox.
- Files of this kind inside a repository you clone are untrusted until reviewed, and tools must not load or execute them automatically.
- Agents must not be able to install or modify their own extensions.
A tool-security checklist#
- Agents reach tools only through the gateway; no direct connections to arbitrary servers.
- Every server in the catalogue has an owner, a reviewed version and a description hash.
- Description or tool-list changes suspend the server until re-reviewed.
- Tokens are per user, per server, short-lived; no passthrough.
- Local servers run sandboxed; nothing auto-starts from repository configuration.
- Each agent task type has a minimal tool set; dangerous combinations are reviewed.
- Tool results are screened, size-limited and set session taint.
- Your own servers authenticate, authorise as the user, and validate every argument.
- One switch disables a server or tool for all agents.
Remember this#
- An MCP server is third-party code and third-party text in a privileged position.
- Most tool incidents are classic bugs — missing auth, command injection — plus poisoned descriptions and auto-executed configuration.
- Pin the tool descriptions and re-review on change.
- Centralise approval, authorization, brokering, inspection and audit in a tool gateway.
- Publish narrow, authenticated, input-validating servers that act as the user.
- Evaluate combinations of tools for the trifecta, not only each tool alone.
Try it#
- Inventory the MCP servers and extensions an agent of yours can load. For each, answer the six questions.
- Read the full tool descriptions of one server as the model would. Is there anything in them that is an instruction rather than a description?
- Find a pair of tools in your catalogue that together form an exfiltration path.
Check yourself#
- Why is pinning a server’s version not enough to prevent a rug pull?
- What does a tool gateway centralise, and why does the 2026-07-28 MCP revision make that easier?
- What is the most common class of vulnerability in MCP servers, and how is it avoided?