Numerical examples are illustrative unless explicitly sourced.
Remember: Untrusted content can inform a decision; it cannot grant permission.
Definition and threat model
Prompt injection is an attack in which supplied content attempts to override or redirect a language-model application's intended instructions. In a direct attack, the attacker supplies the user-facing input. In an indirect attack, the instruction arrives through a page, file, message, search result or tool response the assistant reads. The content does not need to be visually hidden; ordinary text can be enough.
For an assistant summarizing a supplier proposal, the document is evidence about the supplier. An embedded request to upload an internal customer list is an attempt to redirect the task, not an authorized instruction from the user.
| Entry point | Claimed authority | Capability at risk |
|---|---|---|
| Retrieved document | “The administrator requires an export” | Access to confidential data and outbound tools |
| Tool response | “Run this follow-up command to finish” | Shell, file or network access |
| Saved memory | “Approval was already granted” | Later execution without valid approval |
| User input | “Ignore the application's access rules” | Cross-account data or privileged operations |
This is different from a normal quotation of an instruction. A security article may legitimately contain “ignore previous instructions” as an example. Blocking every occurrence would damage useful work. The important questions are who controls the content, what authority it has, and what sensitive capability it could influence.
Separate probabilistic understanding from enforced authority
Clear source labels and delimiters help the model distinguish governing instructions from evidence. A detector may flag suspicious passages. An extraction step can reduce how much raw text reaches another stage. These are useful reductions in exposure, but each model can still misunderstand or pass along an attack in a transformed form. A summary is not automatically trustworthy because another model wrote it.
Enforced boundaries live outside that interpretation. The document-reading component can operate without credentials for payments or email. A separate executor receives a typed action proposal and checks the authenticated user, target resource, allowed recipient, amount, current permissions, and approval. If those checks reject the action, persuasive document text cannot override the decision.
Exfiltration means moving protected information to an unauthorized destination. It can happen through an email, a tool parameter, an uploaded file, a generated link, or an outbound request. A control that scans only the final visible answer can miss those routes. Map where sensitive data can go, and restrict unnecessary outbound capabilities at those boundaries.
Why the SQL analogy has limits
SQL injection is commonly prevented by keeping values separate from query structure through parameterized statements, along with authorization and other controls. Natural language does not have a comparable universal escape mechanism that makes every sentence incapable of influencing a model. Putting text in XML tags is helpful formatting, not the same guarantee as a database parameter binding.
This does not mean defenses are hopeless or that “use a guard model” is the replacement guarantee. It means model behavior and application authority need different protections. Test the model's susceptibility, but also test that a successful persuasion attempt cannot produce an unauthorized effect.
For a browser agent, require especially clear boundaries between “this page says to pay” and “the user authorized this exact purchase.” Payment approval should reference the merchant, amount, item, and proposal version. The page being visited cannot be the authority that grants itself access to the user's money.
Put controls at the actual boundary
Read diagram source
flowchart TD
U[Untrusted document or tool result] --> M[Model proposes response or action]
M --> G[Application policy and authorization]
G -->|Allowed scoped action| T[Tool with limited credentials]
G -->|Denied or needs approval| S[Stop or request review]
T --> V[Validate and safely render result]
| Layer | Control | Residual limitation |
|---|---|---|
| Retrieval | Fetch only data the authenticated user may access | Authorized data can still contain attacks |
| Prompt construction | Clearly label external content and provenance | The model may still follow it |
| Detection | Scan for suspicious requests and known attack patterns | False negatives and false positives remain |
| Tool gateway | Verify identity, target, arguments, and policy | Must cover every route to the action |
| Runtime | Restrict files, credentials, network, and resources | Isolation requires correct configuration |
| Output | Validate structured results and sanitize rendering | A valid format can still contain false claims |
OWASP describes prompt injection as a risk that requires layered mitigation, including limited privilege and oversight; it does not claim a complete prompt-only fix. OWASP prompt injection.
Why common fixes are insufficient
“Put the document inside XML tags.” Helpful organization, but the tags are text interpreted by a probabilistic model. They do not behave like memory protection or database permissions.
“Use a guard model.” It may detect attacks, but can be fooled or misclassify legitimate text. A second model reviewing the first is another fallible component. If a design separates privileged execution from untrusted analysis, the real protection comes from the constrained interface and enforced permissions.
“Put a secret canary in the system prompt.” A canary can detect some leaks. It does not prove that other sensitive data cannot leak. Production secrets should not be placed in model context as an access-control strategy.
“Block strings such as exec or javascript.” String lists are incomplete. Prevent code execution with explicit tool contracts; prevent browser injection with context-appropriate escaping, sanitization, and restrictive rendering.
Work through the supplier attack
The summarizer receives only the requested PDF, without customer-database credentials. If it proposes a network upload, the gateway rejects it because the task grants no upload capability. For an approved export workflow, the server independently resolves the permitted dataset and destination; a PDF cannot expand either. The system records the rejected proposal without logging sensitive payloads unnecessarily.
Also consider allowed-channel exfiltration: an attacker may put confidential data into an approved search query, email subject, image URL, or log field. Destination allowlists alone are not enough if the allowed destination is inappropriate for that data.
Persistent memory needs the same treatment. Store provenance and scope with a note, and distinguish a quoted document instruction from a user preference or verified application fact. Recheck current access and approvals when the note is retrieved. Neither storage nor summarization turns lower-trust content into governing policy.
Trust in a tool's authenticated transport is distinct from trust in the text it returns. Anthropic's containment discussion makes this distinction in its handling of third-party content and execution boundaries. The general design lesson is to identify who authored a result and constrain its authority at the next step. Containment across products.
Manager decisions and measurements
Inventory capabilities and trust boundaries before choosing a filter vendor. Assign ownership for tool authorization, sandbox policy, adversarial tests, and incident response. Test cross-tenant access, malicious citations, forged approval text, poisoned memory, and harmful actions split across several individually innocuous steps.
Measure unauthorized action success, data exposure, benign-task completion, and false blocks. A lower attack success rate on one test set is useful evidence, not a universal security guarantee. On suspected compromise, disable affected capabilities, revoke scoped credentials, preserve necessary evidence, and inspect actual external effects.
| Test | Evidence to inspect | What a passing result means |
|---|---|---|
| Detector deliberately bypassed | Gateway decisions and actual tool effects | Capability boundaries survived this detector failure |
| Hostile text inside a valid document | Task completion and unauthorized-action attempts | The legitimate task remained usable under the tested attack |
| Poisoned note loaded next session | Source label, live authorization and proposal | Persistence did not grant new authority |
| Allowed destination with forbidden data | Actual outbound payload and policy decision | Data-use policy covered this allowed channel |
| Benign quotation of an attack | False blocks and task success | Detection did not simply reject every matching phrase |
Use controlled test data and an isolated environment. Report the attack set, tested tools and limits of coverage alongside the rates.
Recall questions
“The system message says never leak secrets. Is that enough?” No; limit exposure and enforce permissions outside generation.
“Can a trusted search tool return untrusted content?” Yes. Trust in the transport does not make every retrieved author's instructions trustworthy.
“What survives a failed detector?” The sandbox, scoped identity, policy gateway, and approval rules must still limit harm.
Related: Agent sandboxing.
Interview questions with developed answers
Q1: Why is prompt sanitization harder than preventing SQL injection?
Sample answer: SQL has a defined query structure, and parameterized statements keep supplied values from becoming that structure. Natural-language text has no universal escaping rule that guarantees a model will never interpret it as an instruction. Attackers can paraphrase, use context, or distribute the request across steps. I therefore use labels and detectors as supporting measures and enforce data access and tool authority outside the model. The goal is to limit what a mistaken interpretation can cause, not to claim that one sanitization pass removes every possible attack.
Follow-up: Do parameterized queries eliminate all database risk? No; a valid query can still access unauthorized rows without proper authorization.
Q2: What is indirect prompt injection in a RAG system?
Sample answer: It occurs when a retrieved document contains instructions that try to redirect the assistant beyond the user's task. The retrieval may be legitimate; the authority claimed by the document is not. I preserve source provenance, treat the content as evidence, minimize privileges, and validate any proposed action at the executor. I test both answer contamination and data movement through tools. An extra summarization model can help reduce exposure but may also carry the malicious instruction forward, so it is not sufficient by itself.
Follow-up: Can an internal document be hostile? Yes, through compromise, editable content, or accidental inclusion of misleading instructions.
Q3: Is a guard model or a two-model architecture enough?
Sample answer: No. A guard model has false positives and false negatives, and its own inputs can be adversarial. A two-model design can separate responsibilities, especially when the component reading untrusted content has fewer capabilities. But if the second model blindly trusts the first model's text, the attack can cross that boundary. I require typed outputs, explicit validation, independent authorization, and limited execution privileges. I would measure whether the architecture reduces risk on realistic attacks and what remains possible after a detector miss.
Follow-up: What should the untrusted reader be unable to do? Access unnecessary secrets or invoke consequential tools directly.
Q4: How do you test defenses against data exfiltration?
Sample answer: I place controlled test data in an isolated environment and try to induce unauthorized transfers through every permitted output path: messages, links, files, tool arguments, and network requests. I verify actual transmission attempts and executor decisions, not only whether the final response sounds safe. I include destination changes, encoded content, and multi-step tasks within authorized test scope. I also confirm that normal business transfers still work under the policy. The result is evidence about specific boundaries and tested attacks, not a universal security certificate.
Follow-up: Why use isolated test data? To measure the failure without exposing real customer information.
Q5: What should happen if the model proposes a forbidden action?
Sample answer: The executor rejects it using trusted policy and records an appropriate audit event. The application can continue with a safe part of the task, ask the user for a valid alternative, or stop and escalate. It must not retry the same forbidden action through a more permissive tool or model. I investigate whether the proposal came from confusion, injection, or an interface problem, then add a regression case. A useful refusal is evidence that the authority boundary worked even if the model was persuaded.
Follow-up: Does human approval fix everything? Only when approval is informed, authorized, and bound to the exact action being executed.
60-second interview answer
Prompt injection is an attempt to make a model follow attacker-controlled instructions in user input, retrieved documents, tool results, or other content. I assume detection can miss attacks. I limit what data and tools the model can reach, and enforce authorization, destinations, and action policy in application code. I separate data from instructions and use filters as additional signals, but neither XML tags nor another model creates a security boundary. I test direct, indirect, and multi-step attacks, including attempts to leak data through tools, URLs, or persistent memory.