LLM application security protects the confidentiality, integrity and availability of an application that uses a language model. It includes ordinary web, identity, data and infrastructure security, plus risks created when model inputs or outputs influence decisions and actions.
A trust boundary separates components or actors with different authority. A retrieved document may be valid evidence for an answer while having no authority to change permissions or issue commands. The central interview question is where content can cross that boundary.
Define the threats before choosing a filter
| Term | Standard meaning | Concrete failure |
|---|---|---|
| Prompt injection | Attacker-controlled input redirects model behavior against the application's intended instructions | A retrieved page causes an assistant to propose an unauthorized export |
| Direct injection | The attack arrives through the user's interaction | A request tries to override application restrictions |
| Indirect injection | The attack arrives through other content the model processes | A document, image or tool response contains instructions |
| Jailbreak | An attempt to bypass a model's safety restrictions | A prompt elicits prohibited model behavior; it need not involve a connected tool |
| Data poisoning | Deliberate contamination of data used for training or another model pipeline | An attacker inserts misleading records into a training set or retrieval corpus |
| Excessive agency | The application grants more capability, permission or autonomy than its task needs | A summarizer can delete records or send arbitrary external messages |
| Improper output handling | A downstream component consumes generated output without suitable validation or encoding | A browser executes generated markup or a database executes arbitrary generated SQL |
Injection and jailbreak are related terms, not synonyms. A model can leak data without executing a tool, and a completely ordinary API authorization bug can expose data without any prompt injection. See the dedicated prompt-injection lesson and OWASP prevention guidance.
Current risk map
The OWASP GenAI LLM Top 10 2026 was published on August 3, 2026. Its identifiers differ from the 2025 edition still shown on some overview pages. The following short descriptions follow the 2026 document; attach the edition when referring to an ID.
| ID | Risk area | What the design must examine |
|---|---|---|
| LLM01 | Injected instructions | Untrusted content influencing behavior |
| LLM02 | Disclosure of sensitive data | Prompts, answers and secondary copies |
| LLM03 | Overpowered agents | Reachable tools and delegated privileges |
| LLM04 | Compromised supply chain | Packages, models, datasets and tool servers |
| LLM05 | Poisoned data or models | Training, adaptation and evidence integrity |
| LLM06 | Unbounded resource use | Loops, large inputs, spending and capacity exhaustion |
| LLM07 | Misleading generated information | Incorrect claims and unjustified reliance |
| LLM08 | Exposure of hidden context | Internal instructions, memory, tool schemas and other non-user-visible context |
| LLM09 | Weak vector/embedding controls | Retrieval isolation, integrity and leakage |
| LLM10 | Unsafe consumption of output | Renderers, commands and downstream APIs |
A taxonomy helps coverage; it does not prioritize your system's risks automatically. For a feedback tutor, fabricated feedback and private recording exposure matter. For a payment agent, unauthorized transactions add a different consequence.
Interview scope: a document assistant with one write tool
The assistant reads permitted support documents and can create an internal support ticket. It cannot send arbitrary external messages or change payment records.
Functional requirements
- Authenticate the caller and retrieve only currently permitted documents.
- Produce answers with source references and express unsupported conclusions clearly.
- Propose tickets with constrained fields and execute only authorized submissions.
- Retain enough evidence to investigate an access or action failure.
- Let operators disable the affected capability and recover safely.
Non-functional requirements
- Keep secrets outside model context and ordinary logs.
- Prevent cross-tenant reads through retrieval, caches, artifacts and jobs.
- Bound request size, model calls, tool calls, runtime and spending.
- Isolate any generated-code execution and constrain network destinations.
- Measure false blocks and missed attacks on the actual workflow.
Clarify requirements before adding a classifier. A service that does not need external email should not obtain that capability merely because an agent framework exposes it.
Threat model and first design
Record assets, attackers, entry points and consequences. Documents, messages, tool descriptions, tool results, image text, stored memory and package install scripts have different provenance. A trusted transport does not make all content from a server authoritative.
The basic design is an authenticated API, permission-aware retriever, model and a restricted ticket executor. The model proposes content; server code decides whether an operation may run. Preserve trusted identity in application state, separately from model-generated fields.
| Boundary | First control | Remaining failure to address |
|---|---|---|
| Request → application | Verified identity, size and quota limits | Valid users can still submit hostile content |
| Source → retrieval context | Current object permissions and source versions | Permitted documents may contain injection |
| Model → executor | Typed arguments, resource/action authorization | Valid fields can describe an unwanted but permitted action |
| Executor → external service | Scoped credential and target constraints | A timeout may hide a completed write |
| Answer → browser | Text rendering or controlled sanitized markup | Remote images/URLs can disclose information |
| Request → telemetry | Minimized fields and restricted trace access | Debugging can create additional sensitive copies |
For permitted but consequential actions, use the product's explicit delegation or confirmation policy. Confirmation must bind the exact action; repeatedly prompting for already-authorized low-risk work is not a security architecture. Recheck current authority and business state at execution.
Add controls where the design fails
- Context provenance: keep source identifiers and trust metadata in application state. Delimiters and instructions help the model distinguish evidence, but are not a hard execution boundary.
- Detection: screen suspicious content where useful. Keyword patterns miss paraphrases and encodings; model classifiers can also be attacked. Measure them against legitimate inputs as well as attacks.
- Constrained effects: expose narrow operations with tenant, target and amount checks. A sandbox needs resource and network controls as well as filesystem isolation; a container alone is not a complete sandbox.
- Egress policy: validate allowed destinations and redirects; prevent requests to disallowed internal services. Avoid automatically loading model-supplied remote images in a private answer.
- Durable action state: use stable operation IDs, bounded retries and reconciliation for uncertain writes. A refusal message after an action has run does not undo it.
- Supply-chain controls: pin and review executable dependencies, verify artifacts and limit install/runtime credentials. A signed artifact proves something about origin/integrity, not that the code is safe.
Continue with agent sandboxing and access control.
Data protection includes secondary copies
Data can escape through context, output, embeddings, conversation history, traces, exports and backups. Embeddings and unkeyed hashes of predictable content are not automatic anonymization. Classify data, minimize collected content, restrict access and define retention/deletion for each copy.
A private model changes a hosting boundary; it does not fix an overprivileged tool or shared cache. For hosted models, verify the actual service's retention, training-use and processing-location terms and configuration. Do not infer those properties from the vendor name alone.
System instructions are also not a secret vault. Avoid embedding credentials or relying on an unrevealed prompt as the only access policy. OWASP's system-prompt risk guidance emphasizes the underlying exposure of sensitive information and controls.
Compare cost and benefit
| Decision | Benefit | Cost or limitation |
|---|---|---|
| Smaller tool permission set | Reduces reachable harm after a model mistake | More explicit integration work |
| Detector before generation | Blocks some attacks cheaply | False positives, latency and bypasses |
| Extra model verifier | Adds a second behavioral check | Correlated mistakes and extra cost |
| Confirmation for selected actions | Gives the user control over consequences | Friction; users need a clear proposal |
| Restricted raw-trace retention | Reduces exposure and storage | Harder incident reconstruction without selective evidence |
| Provider/model diversity | Limits some shared outages and weaknesses | More contracts and configurations to validate |
A strong close identifies the largest remaining exposure and how to contain it. Do not claim a benchmark score proves injection immunity or that several fallible filters multiply into an independently proven failure probability.
Safe output handling at two different sinks
A sink is the component that consumes output, such as a browser renderer, database or command executor. A generated string is untrusted data even if it passed a content filter. For plain text displayed in a browser, use a text-only API such as textContent; if rendering HTML is required, use an appropriate maintained sanitizer and a restricted rendering policy. HTML escaping alone does not validate URLs, JavaScript, CSS, or every embedding context.
// Text-only rendering: never interpret the model output as markup.
const answerNode = document.createElement("p");
answerNode.textContent = modelOutput;
container.replaceChildren(answerNode);
A separate order-lookup example illustrates the database boundary. Choose the query structure in trusted application code and bind values. This Python/SQLite example is a parameterized data lookup, not permission to execute model-generated SQL:
def permitted_order(db, authenticated_tenant, order_id):
return db.execute(
"SELECT order_id, status FROM orders "
"WHERE tenant_id = ? AND order_id = ?",
(authenticated_tenant, order_id),
).fetchone()
The tenant comes from authenticated application context, never from the model. Parameterization prevents values from changing SQL structure; it does not replace row authorization, read-only database credentials, or restrictions on the fields returned. For an approved analytics SQL feature, parse and constrain the query structure, use least-privilege execution, and test the full policy separately.
Trace authorized retrieval and an authorized tool call
Read diagram source
flowchart TD
U[Authenticated principal and tenant] --> P[Current policy and resource scope]
P --> R[Retrieve only permitted evidence]
R --> M[Model receives evidence as untrusted data]
M --> T[Proposed typed tool arguments]
T --> V[Independent schema, ownership, and business checks]
V --> A{Exact action needs approval}
A -->|Yes| H[Bind approval to action digest and expiry]
H --> E[Recheck and execute with operation ID]
A -->|No| E
E --> O[Validate receipt and minimize returned data]
| Test | Expected result | Boundary exercised |
|---|---|---|
| Tenant A searches its currently permitted policy | Allow; return authorized source revision | Retrieval authorization |
| Tenant A names tenant B's document ID | Deny before content reaches the model | Resource ownership, not prompt compliance |
| A passage instructs the agent to upload all orders | Tool request denied unless independently authorized, which this action is not | Untrusted evidence cannot grant authority |
| A valid JSON ticket changes the approved project | Deny and require a new proposal/approval | Business meaning beyond schema |
| A user loses access while the workflow waits | Deny at resume and exclude stale cached content | Current permission revalidation |
| Ticket creation times out after possible commit | Reconcile the same operation ID | Duplicate-effect prevention |
Run these as integration/adversarial cases across the real cache, retrieval, and tool layers. A test that only asks the model to refuse the forbidden request misses the actual enforcement boundary.
Interview questions with developed answers
Q1: How do you defend against prompt injection?
Sample answer: I trace the path from attacker-controlled content to sensitive data or actions. I label external content as evidence and use detection and model instructions as supporting controls. The decisive boundary is the executor: it checks identity, resource authorization, action parameters, destination, and approval before performing a tool call. I minimize tool privileges and credentials in context, restrict unnecessary outbound access, and test indirect injections in documents and tool results. I also plan for a bypass, so one mistaken model decision has a limited impact.
Follow-up: Why is a stronger system prompt insufficient? It influences model behavior but does not enforce database or tool permissions.
Q2: How do you handle multi-tenant data security in RAG?
Sample answer: Tenant and user scope come from authenticated server state. I enforce access before content leaves the trusted retrieval layer for models, users, shared caches or artifacts. I preserve document-level permissions, recheck revocation, and isolate conversations and asynchronous jobs. I test attempts to alter tenant identifiers and to reuse another user's cached answer. Encryption and output filters are additional controls, but neither compensates for excessive data access. Audit records should show which principal accessed which resource under which policy without exposing unnecessary payloads.
Follow-up: Is a tenant ID enough? Users within the same tenant may have different document permissions.
Q3: Why must generated output be treated as untrusted input?
Sample answer: The model may produce malicious or simply incorrect content, including content influenced by an attacker. The consuming component must enforce its own contract. I validate structured fields, parameterize database values, sanitize display HTML, and restrict action targets and file paths. I avoid turning arbitrary text into a shell command or privileged query. The right control depends on the sink; scanning the answer for suspicious phrases cannot replace these boundaries. I also verify the resulting business state before reporting success.
Follow-up: Does schema validation prevent unauthorized actions? It verifies shape, not permission or business correctness.
Q4: What do you put into a security test plan?
Sample answer: I start from assets and trust boundaries, then test direct and indirect injection, cross-tenant access, permission revocation, malicious tool arguments, sensitive logging, and resource exhaustion. I include multi-step scenarios because a harmful action can arise from several individually plausible inputs. I verify blocked effects at the database, network, or action service rather than relying only on the model's refusal text. Production incidents feed regression cases, and new tools or data sources trigger a review of the threat model.
Follow-up: What does a clean adversarial test prove? Coverage of the tested cases under those conditions, not immunity to all attacks.
Q5: How would you respond to suspected data leakage?
Sample answer: I contain the affected path, preserve relevant evidence under controlled access, and determine the scope of exposed data and users. I revoke compromised credentials where appropriate and involve the organization's security and incident owners. I distinguish a model claiming it saw data from evidence that the data was actually accessible or transmitted. The repair addresses the authorization or data-flow failure, with regression tests and a controlled return to service. Customer and regulatory communication follows the established response process based on verified facts.
Follow-up: Why preserve logs carefully? They may be essential evidence and may themselves contain sensitive data.
60-second interview answer
LLM security extends ordinary application security. I begin with assets, actors, entry points, and trust boundaries, then consider how model-generated text can cross into data access or execution. I enforce authorization outside the model, scope tools and secrets, isolate generated code, validate outputs, and control sensitive logs and caches. Prompt injection detection is useful but fallible. I test the actual system for data exposure and unauthorized effects, including supply-chain and denial-of-service paths, and define containment and recovery before launch.