Today ainotis Join

My notis

SafetyPublished ExplainedAll news from that day

How prompt injection works and how to limit the damage

OWASP, Anthropic and OpenAI describe prompt injection as untrusted text that a model may follow as an instruction, and the guidance centres on limiting what a fooled model can do.

Share

Check our sources · 25 facts from 4 sources
Illustration: A parcel locker with an intake slot and a sealed inner compartment; its green approval latch stands for human approval before a risky action.Illustration made with AI for ai notis
source · Illustration made with AI for ai notis; not a picture of what happened

Key points

  1. OWASP says a prompt injection occurs when prompts alter a model's behaviour in unintended ways, either directly from the user or indirectly through external sources such as websites or files.
  2. OWASP says it is unclear whether fool-proof prevention exists, and that retrieval augmented generation and fine-tuning "do not fully mitigate" prompt injection vulnerabilities.
  3. OWASP lists least privilege, human approval for high-risk actions and separating external content; Anthropic adds running tools in sandboxed environments, and OpenAI says agents can still be tricked.

What happened

OWASP, the open security project, defines a prompt injection vulnerability as one that occurs when user prompts alter the behaviour or output of a language model in unintended ways. The inputs do not need to be readable by a person, as long as the model parses the content.

OWASP separates two kinds. In a direct injection, the user's own prompt changes the model's behaviour. In an indirect injection, the model accepts input from an external source such as a website or a file, and that content changes its behaviour.

Anthropic's documentation draws the same line by who the adversary is: in the first case the user of the application is hostile, and in the second the user is trusted but Claude reads third-party content, such as web pages, emails, documents and tool results, that contains adversarial instructions.

OpenAI's guidance for agent builders describes the goals of an attacker as exfiltrating private data through downstream tool calls, taking misaligned actions, or otherwise changing the model's behaviour. OWASP says the severity depends on the business context and on how much agency the model has been given.

It lists disclosure of sensitive information, unauthorised access to functions available to the model, and execution of arbitrary commands in connected systems among the possible results.

OWASP says it is unclear whether any prevention is fool-proof, given how models work, and that retrieval-augmented generation and fine-tuning do not fully mitigate the problem. Its measures therefore limit the impact. They include giving the model the least privilege it needs, requiring human approval for high-risk actions, and separating and marking untrusted external content.

Anthropic's documentation applies the same ideas. It says to put third-party content only in tool results, to tell Claude what the content is and where it came from, and to state in the system prompt that returned content is untrusted data and must never override the user's request.

It also says to wrap untrusted strings in JSON, to limit Claude's access to sensitive data and actions, and to run tools in sandboxed environments.

OpenAI's guidance adds that untrusted input should not be placed in developer messages, which take precedence over user messages, and that structured outputs between steps remove free-text channels an attacker could use.

Claude Code, Anthropic's coding agent, applies some of this in the product. Its documentation says commands that fetch content from the web, such as curl and wget, are not auto-approved by default, and that sandboxing can add filesystem and network isolation for shell commands. OpenAI's guidance states that even with its mitigations agents will not be perfect and can still be tricked.

What it means for you

Our view

The vendors agree a model cannot be relied on to refuse every injected instruction, so the design question becomes what a fooled model can reach. OWASP ties the severity to the agency the model has been given, and OpenAI says that even with mitigations agents can still make mistakes or be tricked.

For an agent that reads email, web pages or documents, that points to reviewing its permissions as well as its prompt. Ask your team or vendor which tools the agent can call, which need human approval, which run in a sandbox, and whether untrusted content is labelled as data. Start with the one tool that could send data out.

It adds no new facts.

Share this story

Your reaction

We count reactions per story and day, never who reacted. The counts help us choose what goes in the monthly issue. If you are signed in, your own page shows yours too.

Check our sources

Every sentence above is checked against these 4 sources.

1 LLM01:2025 Prompt InjectionOWASP Gen AI Security Project · 11 facts Open the source archived copy
  1. OWASP's LLM01:2025 page says: "A Prompt Injection Vulnerability occurs when user prompts alter the LLM’s behavior or output in unintended ways."

    LLM01:2025 Prompt Injection A Prompt Injection Vulnerability occurs when user prompts alter the LLM's behavior or output in unintended ways.
    cite
  2. OWASP says these inputs "can affect the model even if they are imperceptible to humans" and that prompt injections "do not need to be human-visible/readable, as long as the content is parsed by the model".

    These inputs can affect the model even if they are imperceptible to humans, therefore prompt injections do not need to be human-visible/readable, as long as the content is parsed by the model.
    cite
  3. OWASP says direct prompt injections occur "when a user’s prompt input directly alters the behavior of the model in unintended or unexpected ways".

    Direct prompt injections occur when a user's prompt input directly alters the behavior of the model in unintended or unexpected ways.
    cite
  4. OWASP says indirect prompt injections occur "when an LLM accepts input from external sources, such as websites or files" and that the content, when interpreted by the model, alters its behavior in unintended or unexpected ways.

    Indirect prompt injections occur when an LLM accepts input from external sources, such as websites or files. The content may have in the external content data that when interpreted by the model, alters the behavior of the model in unintended or unexpected ways.
    cite
  5. OWASP says the severity and nature of a successful prompt injection "are largely dependent on both the business context the model operates in, and the agency with which the model is architected".

    are largely dependent on both the business context the model operates in, and the agency with which the model is architected
    cite
  6. OWASP lists "Providing unauthorized access to functions available to the LLM" and "Executing arbitrary commands in connected systems" among the unintended outcomes of prompt injection.

    prompt injection can lead to unintended outcomes, including but not limited to: Disclosure of sensitive information ... Providing unauthorized access to functions available to the LLM ... Executing arbitrary commands in connected systems
    cite
  7. OWASP says: "it is unclear if there are fool-proof methods of prevention for prompt injection".

    it is unclear if there are fool-proof methods of prevention for prompt injection
    cite
  8. OWASP says that techniques like Retrieval Augmented Generation (RAG) and fine-tuning aim to make outputs more relevant and accurate, but "research shows that they do not fully mitigate prompt injection vulnerabilities".

    research shows that they do not fully mitigate prompt injection vulnerabilities
    cite
  9. OWASP's listed mitigations include "Enforce privilege control and least privilege access", "Require human approval for high-risk actions", and "Segregate and identify external content".

    Enforce privilege control and least privilege access ... Require human approval for high-risk actions ... Segregate and identify external content
    cite
  10. OWASP says to "Separate and clearly denote untrusted content to limit its influence on user prompts".

    Separate and clearly denote untrusted content to limit its influence on user prompts.
    cite
  11. OWASP's LLM01:2025 page lists "Disclosure of sensitive information" first among the unintended outcomes prompt injection can lead to.

    prompt injection can lead to unintended outcomes, including but not limited to: Disclosure of sensitive information
    cite
2 Mitigate jailbreaks and prompt injectionsAnthropic (Claude Platform Docs) · 7 facts Open the source archived copy
  1. Anthropic's documentation says jailbreaks and direct prompt injection are the case "where the user of your application is the adversary and crafts inputs intended to bypass your guardrails".

    where the user of your application is the adversary and crafts inputs intended to bypass your guardrails
    cite
  2. Anthropic's documentation says indirect prompt injection is the case "where the user is trusted but Claude processes third-party content (web pages, emails, documents, tool results) that contains adversarial instructions".

    where the user is trusted but Claude processes third-party content (web pages, emails, documents, tool results) that contains adversarial instructions
    cite
  3. Anthropic's documentation says: "Put untrusted content only in tool results." and says Claude is trained to treat instructions that appear inside tool results with appropriate skepticism.

    Put untrusted content only in tool results. Deliver third-party content to Claude inside tool_result blocks, never in system prompts or plain user text blocks. Claude is trained to treat instructions that appear inside tool results with appropriate skepticism.
    cite
  4. Anthropic's documentation says to tell Claude what the content is and where it came from, for example that it is the body of an inbound email from an unknown sender.

    Tell Claude what the content is and where it came from. In the tool's description, or in the structure of the result itself, make the nature and source of the content explicit: for example, that it is the body of an inbound email from an unknown sender
    cite
  5. Anthropic's documentation says to tell Claude explicitly that content returned from tools, documents, or searches is untrusted data and must never override the system prompt or the user's original request.

    State the policy in your system prompt. Tell Claude explicitly that content returned from tools, documents, or searches is untrusted data and must never override the system prompt or the user's original request.
    cite
  6. Anthropic's documentation says: "JSON-encode untrusted content."

    JSON-encode untrusted content. Where possible, wrap third-party strings in a JSON object rather than concatenating them into free-form text.
    cite
  7. Anthropic's documentation says to apply the principle of least privilege, including "run tools in sandboxed environments".

    Apply the principle of least privilege so that a successful injection can do minimal damage: don't give Claude access to secrets it doesn't need, run tools in sandboxed environments, and scope permissions as narrowly as possible.
    cite
3 Safety in building agentsOpenAI (API docs) · 5 facts Open the source archived copy
  1. OpenAI's guide on safety in building agents says: "A prompt injection happens when untrusted text or data enters an AI system, and malicious contents in that text or data attempt to override instructions to the AI."

    A prompt injection happens when untrusted text or data enters an AI system, and malicious contents in that text or data attempt to override instructions to the AI.
    cite
  2. OpenAI's guide says the end goals of prompt injections "can include exfiltrating private data via downstream tool calls, taking misaligned actions, or otherwise changing model behavior in an unintended way".

    can include exfiltrating private data via downstream tool calls, taking misaligned actions, or otherwise changing model behavior in an unintended way
    cite
  3. OpenAI's guide says: "Because developer messages take precedence over user and assistant messages, injecting untrusted input directly into developer messages gives attackers the highest degree of control."

    Because developer messages take precedence over user and assistant messages, injecting untrusted input directly into developer messages gives attackers the highest degree of control.
    cite
  4. OpenAI's guide says that by defining structured outputs between nodes, such as enums and fixed schemas, "you eliminate freeform channels that attackers can exploit to smuggle instructions or data".

    you eliminate freeform channels that attackers can exploit to smuggle instructions or data
    cite
  5. OpenAI's guide says: "even with these mitigations, agents won’t be perfect and can still make mistakes or be tricked".

    However, even with these mitigations, agents won't be perfect and can still make mistakes or be tricked
    cite
4 Security - Claude Code DocsAnthropic (Claude Code Docs) · 2 facts Open the source archived copy
  1. Claude Code's security documentation says: "Commands that fetch content from the web such as curl and wget are not auto-approved by default."

    Commands that fetch content from the web such as curl and wget are not auto-approved by default.
    cite
  2. Claude Code's security documentation lists a sandboxed bash tool that provides "filesystem and network isolation" for bash commands.

    Sandboxed bash tool: Sandbox bash commands with filesystem and network isolation, reducing permission prompts while maintaining security.
    cite

Topics

The morning email

On the mornings we publish: the three top stories and up to four short ones. Free.

We email you a link to confirm. An issue may include one sponsor, always labelled Sponsored · Advertisement. Our emails count opens and clicks, not who made them. Unsubscribe in one click. What we keep