Pidoku

Abuse and Resource Attacks

Basic 40 min Difficulty 2/5 Lesson 04 of 04

Prerequisites Why AI Security Is Different

The idea in one minute#

Not every attack tries to steal data through a hijacked agent. Three other families target the model as a resource. Jailbreaks make it produce what its provider or operator forbids. Extraction copies what is valuable about it — its behaviour, its system prompt, or functional equivalents of its weights. Resource attacks make it burn compute until your service is unavailable or your bill is ruinous, which for token-priced systems has its own name: denial of wallet. All three share a property: the attacker is usually an authenticated user doing nominally allowed things, at volume or with intent you did not plan for.

A picture#

flowchart LR
  ATT[":i-user: <b>Attacker</b><br/><small>often a valid account</small>"]
  ATT -->|"crafted prompts"| JB[":i-skull: <b>Jailbreak</b><br/><small>bypass refusals and rules</small>"]
  ATT -->|"many systematic queries"| EX[":i-search: <b>Extraction</b><br/><small>copy behaviour, prompt, data</small>"]
  ATT -->|"expensive requests, loops"| DW[":i-coins: <b>Denial of wallet</b><br/><small>exhaust budget and capacity</small>"]
  JB --> M[":i-brain: <b>Your model endpoint</b>"]
  EX --> M
  DW --> M
  M --> H1[":i-triangle-alert: <b>Harmful or off-policy output</b><br/><small>liability, reputation</small>"]
  M --> H2[":i-archive: <b>Stolen value</b><br/><small>distilled copy, leaked prompt</small>"]
  M --> H3[":nvidia: <b>GPUs and budget consumed</b><br/><small>outage for real users</small>"]
  class ATT queue
  class JB,EX,DW warn
  class M compute
  class H1,H2,H3 memory

How it really works#

Jailbreaks#

A jailbreak gets a model to act against its safety training or its operator’s instructions. The model is not confused about who is speaking, as in injection; it is persuaded.

FamilyHow it works
Persona and fiction“Write a story in which a character explains…” — the request is wrapped in a frame where refusing seems out of place
Hypothetical and research framingClaims of authorisation, education or testing
ObfuscationEncodings, ciphers, rare languages, deliberate misspellings, splitting a request across turns so no single message is refused
Many-shotFilling a long context with fabricated examples of the model complying
Gradual escalationA conversation that moves in small steps, each acceptable given the last
Optimised suffixesStrings found by automated search that push the model toward compliance; some transfer between models
MultimodalThe request rendered as text in an image, or spoken in audio

What a jailbreak costs you depends on the product. For a general assistant, it is the provider’s content policy at stake. For your application the question is narrower and more useful: what does your system prompt forbid, and what happens when a user gets around it? If the answer is “the model says something off-brand”, that is a quality issue. If it is “the model issues a refund outside policy”, the policy was in the wrong place — rules that matter must be enforced in code.

Jailbreak resistance has improved considerably with newer models and with dedicated classifiers on input and output. It remains a rate, measured by red teaming, not a guarantee.

Extraction#

System-prompt extraction. With enough attempts a user can usually recover a system prompt, verbatim or in substance. Treat prompts as published. Their value should be in the engineering around them; their content must not include secrets.

Model distillation. By sending many queries and recording the answers, a competitor can train a cheaper model that imitates yours — particularly effective against fine-tuned models whose value lies in a narrow task. The defence is economic and behavioural: per-account limits, detection of systematic querying, terms of service, and not exposing more than the product needs — raw log-probabilities, for instance, make copying far easier.

Training-data extraction. Models can reproduce passages of their training data, more so when the data was repeated or the model was fine-tuned on a small private set. If private records went into fine-tuning, assume some can come out.

Membership inference. Determining whether a specific record was in the training set — which can itself be a privacy breach.

Weight theft. The direct route: stealing the model file from storage, a registry or a serving node. This is an infrastructure security problem and is prevented the usual way — access control, encryption, egress monitoring on very large transfers — plus, for the highest sensitivity, confidential computing.

Resource attacks#

A language model is the most expensive component most systems have ever exposed to the internet, and its cost per request varies by four orders of magnitude.

AttackMechanism
Token floodingMaximum-length inputs and outputs, at whatever request rate the limit allows
Context stuffingHuge uploads or pasted documents that fill the window on every turn
Output amplificationPrompts engineered for the longest possible response
Reasoning amplificationInputs that make a reasoning model think for far longer than usual
Agent loopsTasks, or injected content, that keep an agent calling tools indefinitely
Fan-outOne request that spawns many sub-agents or parallel tool calls
Retrieval and tool abuseDriving expensive downstream services through the model
Free-tier and trial abuseMany accounts, scripted, reselling your capacity
Stolen keysA leaked API key used for someone else’s workload — a common and very expensive incident

On hosted APIs the result is a bill; on your own GPUs it is saturation and an outage for legitimate users. Both are the same attack.

The controls are the limits from system design, which are security controls here:

per request   input cap, output cap (max_tokens), reasoning cap, timeout
per task      step, token, time and spend budgets; fan-out limit; loop detection
per identity  tokens per minute, concurrent requests, spend per day
per tenant    hard budget that stops traffic, not only an alert
globally      admission control and load shedding; anomaly alerts on spend
for keys      short-lived, scoped, never in client code or repositories; rotation; leak scanning

A request-per-minute limit alone does nothing useful: one request can be a thousand times the cost of another.

Misuse of capabilities#

The last abuse family is using your system as a tool for something else: generating spam or phishing at scale, automating fraud, producing content you would not want attributed to you, or using an agent with a browser or a shell as a launch point for attacks on third parties. Controls are a mixture of the above plus usage policy, account verification proportional to capability, monitoring for abuse patterns, and — for agents that can act on external systems — the outbound restrictions already covered.

Telling abuse from use#

The difficulty is that heavy legitimate use and abuse look alike. Signals that help:

  • Shape, not volume: near-identical prompts with systematic variation; maximum-length everything; activity at machine regularity.
  • Economics: an account whose cost to you far exceeds what it pays.
  • Outcome signals: a rising refusal rate, classifier hits, repeated policy denials.
  • Novelty: a sudden change in an account’s model, token or tool-use pattern.

Respond in steps — slow down, require verification, cap, suspend — rather than only with a block, since false positives here are your best customers.

Remember this#

  • Jailbreaks persuade; injection confuses. Rules that matter are enforced in code.
  • Assume system prompts become public; keep secrets out of them.
  • Extraction is fought with limits, monitoring and not exposing more than needed.
  • Token-priced systems are open to denial of wallet; limit tokens and spend, not requests.
  • A leaked API key is the simplest and costliest resource attack.

Try it#

  1. Estimate the maximum cost of one request to an endpoint you run, and the maximum cost per hour for one account at its current limits. Is the number acceptable?
  2. Write the budgets for an agent endpoint using the six-line template above.
  3. Read your product’s system prompt as if it will be published tomorrow. What would you remove?

Check yourself#

  1. How does a jailbreak differ from a prompt injection?
  2. Why is a request-rate limit ineffective against denial of wallet?
  3. Which extraction risks follow from fine-tuning on private data?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom