- OWASP says a prompt injection occurs when prompts alter a model's behaviour in unintended ways, either directly from the user or indirectly through external sources such as websites or files.
- OWASP says it is unclear whether fool-proof prevention exists, and that retrieval augmented generation and fine-tuning "do not fully mitigate" prompt injection vulnerabilities.
- OWASP lists least privilege, human approval for high-risk actions and separating external content; Anthropic adds running tools in sandboxed environments, and OpenAI says agents can still be tricked.
OWASP, the open security project, defines a prompt injection vulnerability as one that occurs when user prompts alter the behaviour or output of a language model in unintended ways. The inputs do not need to be readable by a person, as long as the model parses the content.
OWASP separates two kinds. In a direct injection, the user's own prompt changes the model's behaviour. In an indirect injection, the model accepts input from an external source such as a website or a file, and that content changes its behaviour. Anthropic's documentation draws the same line by who the adversary is: in the first case the user of the application is hostile, and in the second the user is trusted but Claude reads third-party content, such as web pages, emails, documents and tool results, that contains adversarial instructions.
OpenAI's guidance for agent builders describes the goals of an attacker as exfiltrating private data through downstream tool calls, taking misaligned actions, or otherwise changing the model's behaviour. OWASP says the severity depends on the business context and on how much agency the model has been given. It lists disclosure of sensitive information, unauthorised access to functions available to the model, and execution of arbitrary commands in connected systems among the possible results.
OWASP says it is unclear whether any prevention is fool-proof, given how models work, and that retrieval-augmented generation and fine-tuning do not fully mitigate the problem. Its measures therefore limit the impact. They include giving the model the least privilege it needs, requiring human approval for high-risk actions, and separating and marking untrusted external content.
Anthropic's documentation applies the same ideas. It says to put third-party content only in tool results, to tell Claude what the content is and where it came from, and to state in the system prompt that returned content is untrusted data and must never override the user's request. It also says to wrap untrusted strings in JSON, to limit Claude's access to sensitive data and actions, and to run tools in sandboxed environments. OpenAI's guidance adds that untrusted input should not be placed in developer messages, which take precedence over user messages, and that structured outputs between steps remove free-text channels an attacker could use.
Claude Code, Anthropic's coding agent, applies some of this in the product. Its documentation says commands that fetch content from the web, such as curl and wget, are not auto-approved by default, and that sandboxing can add filesystem and network isolation for shell commands. OpenAI's guidance states that even with its mitigations agents will not be perfect and can still be tricked.
What it means for you Our view
The vendors agree a model cannot be relied on to refuse every injected instruction, so the design question becomes what a fooled model can reach. OWASP ties the severity to the agency the model has been given, and OpenAI says that even with mitigations agents can still make mistakes or be tricked. For an agent that reads email, web pages or documents, that points to reviewing its permissions as well as its prompt. Ask your team or vendor which tools the agent can call, which need human approval, which run in a sandbox, and whether untrusted content is labelled as data. Start with the one tool that could send data out.
TAGSAnthropicOWASPOpenAIPrompt injectionSandboxing