Pidoku

Why AI Security Is Different

Foundations 35 min Difficulty 1/5 Lesson 01 of 03

The idea in one minute#

Classic software keeps code and data apart: the program is fixed, and input is something it processes. A language model has no such separation. Its “program” for this request is the text in its context window, and the document you asked it to summarise sits in that same window. If the document says “ignore the above and do this instead”, the model may comply — not because of a bug to be patched, but because following instructions in text is what it does.

Three consequences define the field: any text the model reads can become an instruction, no filter reliably tells the difference, and so safety has to come from what the system around the model permits, not from what the model is told.

A picture#

flowchart TB
  subgraph WEB["A classic web application"]
    direction LR
    I1[":i-user: Input"] --> V1[":i-shield-check: <b>Parameterised query</b><br/><small>data can never become code</small>"] --> DB1[(":postgresql: Database")]
  end
  subgraph LLM["An LLM application"]
    direction LR
    SP[":i-code: <b>Your instructions</b>"] --> CW
    I2[":i-user: <b>User text</b>"] --> CW
    DOC[":i-globe: <b>Document, web page,<br/>tool result</b>"] --> CW
    CW[":i-layers: <b>One context window</b><br/><small>all of it is just tokens</small>"] --> M[":i-brain: <b>Model</b>"]
    M --> ACT[":i-wrench: <b>Actions</b>"]
  end
  class I1,I2 neutral
  class V1 queue
  class DB1 memory
  class SP neutral
  class DOC warn
  class CW io
  class M compute
  class ACT warn

SQL injection was solved by giving data a separate channel that the database never interprets as code. There is no equivalent for a language model. That is the whole difference.

How it really works#

What is new#

PropertyWhy it matters for security
Instructions and data share one channelText from an attacker can redirect the model: prompt injection
Behaviour is learned, not writtenYou cannot read the code to prove what it will do; you can only test, and tests are samples
Non-determinismAn attack that fails nine times may succeed on the tenth; “we tried it and it refused” proves little
Natural-language attack surfaceAttacks need no exploit code; they can be written by anyone, in any language, hidden in images or documents
The model holds what it was shownAnything in the context — and sometimes the training data — can be coaxed back out
Models actWith tools, a manipulated model does not just say something wrong; it does something wrong
A new supply chainWeights, datasets, prompts, tools and MCP servers are dependencies that existing scanners do not understand

The confused deputy#

The oldest name for the central problem is the confused deputy: a program with authority is tricked into using that authority on behalf of someone who lacks it. An AI assistant acting for Alice reads an email from Mallory. The email contains instructions. The assistant, holding Alice’s permissions, carries them out. Mallory never authenticated, never exploited memory corruption, never touched Alice’s account — she sent an email.

Every serious AI security incident has this shape: an attacker’s text, the victim’s authority.

Why there is no clean fix#

It is tempting to believe the next model, or a better filter, will solve injection. The evidence so far says otherwise:

  • Defences that detect or resist injection are probabilistic. A research group testing twelve published defences with adaptive attacks — attackers who study the defence and adjust — bypassed most of them more than 90% of the time.
  • A defence that blocks 99% of attempts is a strong control against accidents and a weak one against an adversary, who simply tries a hundred variations.
  • The space of possible inputs is all of natural language, plus images and audio. It cannot be enumerated.

Newer models are measurably harder to manipulate, and guardrail models help. The engineering conclusion is still the one OWASP states plainly in its 2026 guidance: assume the model will be fooled, and design so that when it is, little can happen.

What is not new#

Most of an AI system is ordinary software, and most breaches of AI systems are ordinary breaches: leaked API keys, unauthenticated endpoints, vulnerable dependencies, over-broad cloud permissions, a debug interface exposed to the internet. Scanners in 2026 still find hundreds of MCP servers online with no authentication at all. Classic security practice applies in full — authentication, authorisation, least privilege, patching, logging, segmentation — and neglecting it because “AI security is about prompts” is the most common mistake of all.

The two questions#

Everything in this course reduces to two questions about a design:

  1. What can influence the model? — every source of text, image or audio reaching its context, and who controls each.
  2. What can the model influence? — every effect its output can cause.

An AI system is secure when every path from an attacker-controlled answer to question 1 to a damaging answer to question 2 is blocked by a control that does not depend on the model behaving. The rest is detail — important detail, which the following topics provide.

A vocabulary to start with#

TermMeaning
Prompt injectionAttacker-supplied text that changes what the model does
Direct injectionThe attacker is the user, typing it
Indirect injectionThe text arrives through content the system fetched: a page, a file, an email, a tool result
JailbreakGetting a model to violate its own safety training or its operator’s rules
ExfiltrationGetting data out to somewhere the attacker can read it
PoisoningCorrupting what the model learns from or retrieves
Excessive agencyA model holding more capability or permission than its task needs
GuardrailA check on input or output, usually a classifier or a second model

Remember this#

  • A model has one channel for instructions and data. That is the root of AI-specific risk.
  • The typical attack is a confused deputy: the attacker’s text, the victim’s authority.
  • Detection is probabilistic; assume the model will sometimes be fooled.
  • Safety comes from limiting what the model’s output can cause.
  • Ordinary security failures still cause most incidents.

Try it#

  1. For an AI feature you use, answer the two questions: what can influence the model, and what can it influence?
  2. Find one place in that product where text written by a stranger reaches the model.
  3. List three classic security controls an AI application needs that have nothing to do with prompts.

Check yourself#

  1. Why did parameterised queries solve SQL injection, and why is there no equivalent for prompt injection?
  2. What is a confused deputy?
  3. Why is a 99%-effective injection filter insufficient against a determined attacker?

Sources#

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom