The idea in one minute#
Not every attack tries to steal data through a hijacked agent. Three other families target the model as a resource. Jailbreaks make it produce what its provider or operator forbids. Extraction copies what is valuable about it — its behaviour, its system prompt, or functional equivalents of its weights. Resource attacks make it burn compute until your service is unavailable or your bill is ruinous, which for token-priced systems has its own name: denial of wallet. All three share a property: the attacker is usually an authenticated user doing nominally allowed things, at volume or with intent you did not plan for.
A picture#
flowchart LR ATT[":i-user: <b>Attacker</b><br/><small>often a valid account</small>"] ATT -->|"crafted prompts"| JB[":i-skull: <b>Jailbreak</b><br/><small>bypass refusals and rules</small>"] ATT -->|"many systematic queries"| EX[":i-search: <b>Extraction</b><br/><small>copy behaviour, prompt, data</small>"] ATT -->|"expensive requests, loops"| DW[":i-coins: <b>Denial of wallet</b><br/><small>exhaust budget and capacity</small>"] JB --> M[":i-brain: <b>Your model endpoint</b>"] EX --> M DW --> M M --> H1[":i-triangle-alert: <b>Harmful or off-policy output</b><br/><small>liability, reputation</small>"] M --> H2[":i-archive: <b>Stolen value</b><br/><small>distilled copy, leaked prompt</small>"] M --> H3[":nvidia: <b>GPUs and budget consumed</b><br/><small>outage for real users</small>"] class ATT queue class JB,EX,DW warn class M compute class H1,H2,H3 memory
How it really works#
Jailbreaks#
A jailbreak gets a model to act against its safety training or its operator’s instructions. The model is not confused about who is speaking, as in injection; it is persuaded.
| Family | How it works |
|---|---|
| Persona and fiction | “Write a story in which a character explains…” — the request is wrapped in a frame where refusing seems out of place |
| Hypothetical and research framing | Claims of authorisation, education or testing |
| Obfuscation | Encodings, ciphers, rare languages, deliberate misspellings, splitting a request across turns so no single message is refused |
| Many-shot | Filling a long context with fabricated examples of the model complying |
| Gradual escalation | A conversation that moves in small steps, each acceptable given the last |
| Optimised suffixes | Strings found by automated search that push the model toward compliance; some transfer between models |
| Multimodal | The request rendered as text in an image, or spoken in audio |
What a jailbreak costs you depends on the product. For a general assistant, it is the provider’s content policy at stake. For your application the question is narrower and more useful: what does your system prompt forbid, and what happens when a user gets around it? If the answer is “the model says something off-brand”, that is a quality issue. If it is “the model issues a refund outside policy”, the policy was in the wrong place — rules that matter must be enforced in code.
Jailbreak resistance has improved considerably with newer models and with dedicated classifiers on input and output. It remains a rate, measured by red teaming, not a guarantee.
Extraction#
System-prompt extraction. With enough attempts a user can usually recover a system prompt, verbatim or in substance. Treat prompts as published. Their value should be in the engineering around them; their content must not include secrets.
Model distillation. By sending many queries and recording the answers, a competitor can train a cheaper model that imitates yours — particularly effective against fine-tuned models whose value lies in a narrow task. The defence is economic and behavioural: per-account limits, detection of systematic querying, terms of service, and not exposing more than the product needs — raw log-probabilities, for instance, make copying far easier.
Training-data extraction. Models can reproduce passages of their training data, more so when the data was repeated or the model was fine-tuned on a small private set. If private records went into fine-tuning, assume some can come out.
Membership inference. Determining whether a specific record was in the training set — which can itself be a privacy breach.
Weight theft. The direct route: stealing the model file from storage, a registry or a serving node. This is an infrastructure security problem and is prevented the usual way — access control, encryption, egress monitoring on very large transfers — plus, for the highest sensitivity, confidential computing.
Resource attacks#
A language model is the most expensive component most systems have ever exposed to the internet, and its cost per request varies by four orders of magnitude.
| Attack | Mechanism |
|---|---|
| Token flooding | Maximum-length inputs and outputs, at whatever request rate the limit allows |
| Context stuffing | Huge uploads or pasted documents that fill the window on every turn |
| Output amplification | Prompts engineered for the longest possible response |
| Reasoning amplification | Inputs that make a reasoning model think for far longer than usual |
| Agent loops | Tasks, or injected content, that keep an agent calling tools indefinitely |
| Fan-out | One request that spawns many sub-agents or parallel tool calls |
| Retrieval and tool abuse | Driving expensive downstream services through the model |
| Free-tier and trial abuse | Many accounts, scripted, reselling your capacity |
| Stolen keys | A leaked API key used for someone else’s workload — a common and very expensive incident |
On hosted APIs the result is a bill; on your own GPUs it is saturation and an outage for legitimate users. Both are the same attack.
The controls are the limits from system design, which are security controls here:
per request input cap, output cap (max_tokens), reasoning cap, timeout
per task step, token, time and spend budgets; fan-out limit; loop detection
per identity tokens per minute, concurrent requests, spend per day
per tenant hard budget that stops traffic, not only an alert
globally admission control and load shedding; anomaly alerts on spend
for keys short-lived, scoped, never in client code or repositories; rotation; leak scanningA request-per-minute limit alone does nothing useful: one request can be a thousand times the cost of another.
Misuse of capabilities#
The last abuse family is using your system as a tool for something else: generating spam or phishing at scale, automating fraud, producing content you would not want attributed to you, or using an agent with a browser or a shell as a launch point for attacks on third parties. Controls are a mixture of the above plus usage policy, account verification proportional to capability, monitoring for abuse patterns, and — for agents that can act on external systems — the outbound restrictions already covered.
Telling abuse from use#
The difficulty is that heavy legitimate use and abuse look alike. Signals that help:
- Shape, not volume: near-identical prompts with systematic variation; maximum-length everything; activity at machine regularity.
- Economics: an account whose cost to you far exceeds what it pays.
- Outcome signals: a rising refusal rate, classifier hits, repeated policy denials.
- Novelty: a sudden change in an account’s model, token or tool-use pattern.
Respond in steps — slow down, require verification, cap, suspend — rather than only with a block, since false positives here are your best customers.
Remember this#
- Jailbreaks persuade; injection confuses. Rules that matter are enforced in code.
- Assume system prompts become public; keep secrets out of them.
- Extraction is fought with limits, monitoring and not exposing more than needed.
- Token-priced systems are open to denial of wallet; limit tokens and spend, not requests.
- A leaked API key is the simplest and costliest resource attack.
Try it#
- Estimate the maximum cost of one request to an endpoint you run, and the maximum cost per hour for one account at its current limits. Is the number acceptable?
- Write the budgets for an agent endpoint using the six-line template above.
- Read your product’s system prompt as if it will be published tomorrow. What would you remove?
Check yourself#
- How does a jailbreak differ from a prompt injection?
- Why is a request-rate limit ineffective against denial of wallet?
- Which extraction risks follow from fine-tuning on private data?