Prompt Injection Defense
Prompt injection defense protects an LLM application when untrusted content tries to redirect its behavior. It focuses on preserving the boundary between instructions and data, while limiting the information and actions an attacker could obtain even if the model follows a malicious instruction.
What it is
An LLM can encounter instructions inside a user's request or indirectly inside documents, web pages and tool results. The attacker attempts to make those instructions override the application's intended task. Unlike a conventional parser, the model does not reliably enforce a formal separation between executable instructions and ordinary language. Defense therefore combines contextual guidance with controls outside the model: trusted identity, restricted tools, validated arguments and constrained destinations. The relevant threat is an unauthorized effect, such as disclosure or a transaction, rather than merely unusual wording in an answer. Retrieval does not make third-party instructions trustworthy.
What the work involves
A practitioner maps every untrusted input and identifies what the model can read, send or change after consuming it. They constrain tools to narrow operations, enforce authorization in application code and require appropriate review for consequential actions. Tests should place attacks in realistic retrieved material, metadata and tool responses, including multistep conversations. A useful artifact connects each attack path to an enforceable boundary and a regression case. Input detectors can add coverage, but their failure must not grant the model new privileges or unrestricted network access.
Illustrative example
An assistant reads a supplier document that tells it to upload the current customer list to a diagnostic website. The application treats the document as evidence, not authorization. Its upload tool accepts only approved destinations and permitted data, so the requested transfer fails even if the model attempts it. The team records the attempt and adds the document variant to tests covering both tool calls and generated answers.
Limits and common mistakes
No prompt wording or injection classifier guarantees that the model will ignore every attack. Overly aggressive filters can block legitimate quoted instructions, while attackers can exploit encodings or longer contexts. A defense should be evaluated through prevented disclosure and actions, alongside false rejections. Separating roles in the prompt helps communication, but genuine security boundaries come from the application's identity, permission and execution controls.
Prerequisites
Defending against prompt injection requires understanding how prompts work — attackers exploit the same mechanisms that prompt engineers use
Related skills
- → is part of: AI Data Security
Sources and further reading
- OWASP: prompt injection
Supports direct and indirect prompt-injection threats and layered mitigations.
Last updated: 2026-10-10