A guardrail is a check or constraint intended to keep an AI application within defined behavior or operating limits. It may be a deterministic rule, a statistical detector or a workflow control. It does not mean the whole system is guaranteed safe.
Examples include a maximum input size, a schema validator, a content classifier, an authorized tool allowlist and a spending limit. Each has a different contract. A classifier deciding whether text is abusive cannot replace a database permission check.
Define the rule and failure behavior
| Control type | Example | What a passing result establishes |
|---|---|---|
| Syntax/schema | Integer cents, required fields, no unknown fields | The proposal has the required structure |
| Business invariant | Refund no greater than remaining balance | The checked state permits that amount; concurrency still matters |
| Authorization | Caller may refund this particular order | That action is permitted under the checked policy |
| Statistical detector | Harmful-content or PII classification | The detector did not flag it under this model/threshold |
| Grounding check | Claim linked to supporting evidence | Measured evidence support, subject to checker/source errors |
| Operational limit | Maximum steps, time or tokens | The application has bounded that resource dimension |
A false positive flags an acceptable case. A false negative misses a case that should be flagged. Here, “positive” means the detector flags a violation; make that convention explicit before discussing rates.
Fail closed means withholding the protected operation when a required check cannot establish permission or acceptability. Fail open means allowing the operation despite that check being unavailable. Choose separately for each control: an optional style check can have a different outage policy from payment authorization.
Interview scope: customer support with refunds
Functional requirements
- Answer permitted support questions using current policy evidence.
- Recognize missing information and offer clarification or escalation.
- Validate refund proposals and execute only eligible, authorized operations.
- Communicate a verified outcome, including pending or failed outcomes.
- Record policy decisions and let operators investigate and appeal false blocks.
Non-functional requirements
- Prevent unauthorized data access and effects independently of model behavior.
- Bound repair attempts, tool calls, latency and cost.
- Specify which content may be shown before validation finishes.
- Measure missed violations, legitimate work blocked and review-queue load.
- Define degraded behavior when a checker or provider is unavailable.
Start with authentication, authorized retrieval, one model, deterministic proposal checks and a controlled executor. Add content or grounding detectors for identified failure modes. “Add every available filter” increases latency and false blocks without demonstrating protection.
Place checks where they can prevent harm
| Boundary | Check | Failure response |
|---|---|---|
| Input | Format, size, unsupported attachment, malware where relevant | Reject, clarify or quarantine |
| Evidence | Current ACL, source version and relevance | Exclude unauthorized material; clarify insufficient evidence |
| Proposal | Schema, permitted operation, target and business conditions | Repair format once or deny the operation |
| Execution | Current authorization, required approval, balance and operation ID | Stop, request needed authorization or reconcile |
| User output | Content, disclosure, key claims and safe rendering | Withhold, use verified template or hand off |
| Runtime | Deadline, call budget, concurrency and spend | Cancel bounded work and report status |
Do not remove all identifiers indiscriminately: a support workflow may legitimately need an order ID. Minimize and protect necessary data. Regexes can recognize some formats but do not identify all personal information in every language or context.
Schema-constrained generation improves structure. Access control and transactional business logic decide what may happen. A model-generated confirmed: true field is not a user's approval.
Work the confusion matrix
Assume a labeled sample of 10,000 messages, with 100 actual violations. The detector catches 90 and flags 198 acceptable messages. These are illustrative measurements, not a product benchmark.
| Detector decision | Actual violation | Actually acceptable | Total |
|---|---|---|---|
| Flag | 90 true positives | 198 false positives | 288 |
| Allow | 10 false negatives | 9,702 true negatives | 9,712 |
| Total | 100 | 9,900 | 10,000 |
- Recall: 90 / 100 = 90% of violations caught.
- False-positive rate: 198 / 9,900 = 2% of acceptable cases flagged.
- Precision: 90 / 288 = 31.25% of flags are real violations.
- Accuracy: (90 + 9,702) / 10,000 = 97.92%.
Allowing everything would achieve 99% accuracy on this imbalanced set while missing every violation. Accuracy alone is therefore a poor operating target. If all 288 flagged cases require two minutes of review, the queue needs 576 minutes, or 9.6 reviewer-hours, per 10,000 messages.
Choose thresholds using consequence, prevalence, language/task slices and available review capacity. A threshold calibrated on a balanced test set may have very different precision in production. Track uncertain labels and appeal outcomes rather than treating the grader as unquestionable truth.
Streaming changes the enforcement point
A check after emission cannot make text unseen. Choose deliberately:
| Strategy | Benefit | Limitation |
|---|---|---|
| Validate the complete answer before showing it | Checker sees full context before exposure | Delays first content; still subject to checker errors |
| Validate buffered segments with overlap | Earlier useful output | Cross-segment meaning and later contradictions may be missed |
| Emit immediately, then inspect | Lowest perceived delay | Detection is monitoring/containment after partial exposure |
| Stream progress; hold the consequential answer/action | Keeps users informed while checks run | Progress messages must not imply unverified success |
Current NeMo Guardrails output-streaming configuration documents stream_first: true as its default: tokens reach the client before the chunk's output checks. Use stream_first: false when that chunk must be checked first, and test the exact integration. This does not make chunk checks equivalent to full-answer validation. NVIDIA streaming configuration.
Grounding, truth and relevance are different checks
An answer may match the question's topic yet be false. Embedding similarity is a relevance signal, not a factuality verdict. A claim can be supported by an obsolete policy but wrong for today's transaction. Verify source applicability and effective date as well as support.
A grounding checker should distinguish supported, contradicted, insufficient evidence and checker error. Never implement “anything that does not start with NO passes”: empty output, refusal, parse failure and network error would all bypass the check. A citation's existence is not proof that the source supports the claim.
Repeated model agreement also does not establish truth. Multiple samples may repeat the same wrong assumption. Use verifiable evidence and testable constraints before paying for ensembles.
Current implementation choices
| Option | Useful responsibility | Integration caveat |
|---|---|---|
| JSON Schema / Pydantic | Typed data contracts and local checks | Strictness and JSON/Python coercion behavior need tests |
| NeMo Guardrails | Rails around input, dialogue, retrieval, execution and output | Check streaming order and which paths actually invoke each rail |
| Guardrails AI | Composable validators and configured failure actions | Current validators install as separate PyPI packages; select explicit failure behavior |
| OpenAI moderation | Classify supported text/image content categories | A moderation result is not resource authorization or factual verification |
| Application policy and action service | Ownership, eligibility, delegation and idempotency | Must cover workers, retries and alternate entry points |
Current Guardrails AI documentation uses validator packages imported under guardrails_ai, rather than assuming every validator exists in guardrails.validators. Its in-code installation SDK is deprecated. Prefer pinned build-time dependencies. Validator documentation.
The current OpenAI standalone SDK interface is client.moderations.create(...) with a supported moderation model, such as omni-moderation-latest. The current guide also documents moderation integrated with generation. In either mode, handle errors explicitly and test when results arrive relative to streaming or tool execution. Moderation guide. Framework-specific examples here are reference-checked; no hosted classifier is required to run the local validation exercise below.
Failure policies and operating cost
| Failure | Useful response | Unhelpful response |
|---|---|---|
| Malformed proposal | Bounded format repair; rerun all checks | Infinite regeneration |
| Missing evidence | Clarify, abstain or route to an accountable reviewer | Fabricate a source |
| Permission denied | Stop the protected operation | Try providers until one agrees |
| Required checker unavailable | Withhold the affected capability or use a permitted degraded path | Treat an exception as a pass |
| Tool result uncertain | Reconcile by operation ID; show pending status | Retry as a new action and claim success |
| Benign content falsely blocked | Clear explanation and review/appeal route | Silently discard the user's work |
Assign a policy owner and measure cost per resolved task, added latency, false blocks, missed severe failures and human-review time. Privacy and safety detectors can correlate; multiplying their individual miss rates is unjustified without the corresponding independence assumptions. Reassess after model, policy and traffic changes.
Schema, business rules, and bounded repair in one example
Suppose the model proposes a USD refund. A schema can require fields, integer cents, and an allowed currency; it cannot establish that the order belongs to the user or that money remains refundable. Keep those checks against authoritative state.
{
"type": "object",
"additionalProperties": false,
"required": ["order_id", "amount_cents", "currency"],
"properties": {
"order_id": {"type": "string", "minLength": 1},
"amount_cents": {"type": "integer", "minimum": 1, "maximum": 10000},
"currency": {"type": "string", "enum": ["USD"]}
}
}
A maintained validation-library integration can enforce the corresponding application contract. This example targets Pydantic v2; pin and test the dependency. It validates structure and selected business rules only and makes no payment. Its tenant check does not establish that a particular user owns the order or may issue a refund; the executor must separately authorize that operation and check the current order status.
from typing import Literal
from pydantic import BaseModel, ConfigDict, Field
class RefundProposal(BaseModel):
model_config = ConfigDict(extra="forbid", strict=True)
order_id: str = Field(min_length=1)
amount_cents: int = Field(gt=0, le=10000)
currency: Literal["USD"]
def validate_business(proposal, order, principal_tenant):
if order["tenant_id"] != principal_tenant:
raise PermissionError("Wrong tenant")
if proposal.order_id != order["id"] or proposal.currency != order["currency"]:
raise ValueError("Order or currency mismatch")
if proposal.amount_cents > order["refundable_cents"]:
raise ValueError("Amount exceeds current refundable balance")
return proposal
See Pydantic validation and validators for the library contract. JSON validation is one guardrail layer; current policy, approval, concurrency control, and receiver-side deduplication still belong at the action boundary.
| Proposed value or condition | Outcome |
|---|---|
| 4,000 integer cents, correct order/tenant/currency, balance 5,000 | Valid proposal; still check policy and required approval before execution |
"4000" as a string under this strict application contract |
Reject type mismatch |
| 6,000 cents with current balance 5,000 | Schema may pass, business check fails |
| 4,000 cents for another tenant's order | Deny; do not repair by changing tenant identity |
Unknown field such as skip_approval: true |
Reject rather than treating model text as a policy override |
Read diagram source
flowchart TD
M[Model proposal] --> S{Schema valid}
S -->|No, first attempt| R[One bounded format repair]
R --> S2{Schema valid after repair}
S2 -->|No| F[Stop or hand off with reason]
S2 -->|Yes| B[Business and permission checks]
S -->|Yes| B
B --> A{Authorized and approved}
A -->|No| F
A -->|Yes| E[Execute through controlled tool]
Allow at most one format-repair call in this example, within the original deadline and token budget. Return machine-readable validation errors without leaking secrets. Re-run all checks on the repaired output. Never “repair” an authorization denial into an allowed action, and never let repeated repairs become an unbounded agent loop. If the checker is unavailable, a financial write fails closed or hands off; a low-risk text feature may have a separately approved degraded mode.
Interview questions with developed answers
Q1: How do you prevent hallucination in a production RAG system?
Sample answer: I would describe reduction and containment rather than promise complete prevention. I improve source quality and retrieval, preserve necessary qualifications, instruct the model to use evidence, and check important claims and citations. I test unanswerable and conflicting-source cases so clarification and abstention are meaningful. A lower temperature or agreement across samples does not prove truth. For consequential claims I use stronger verification or review, and I measure failures in production. A source-supported answer can still be wrong if the source is obsolete or inapplicable.
Follow-up: What if the answer is unsupported but factually true? Whether it is acceptable depends on the product's evidence contract; distinguish support from truth.
Q2: How do you protect an LLM application from prompt injection?
Sample answer: I assume that external text may contain instructions the application should not follow. Delimiters, model instructions, and classifiers help, but I enforce permissions and action policy in server code. I limit available tools, destinations, and credentials, validate proposed arguments, and bind any approval to the exact action. I test indirect injections in retrieved documents and tool results, including multi-step attacks. A second model can also be manipulated, so it is a supporting detector rather than the sole authority for a sensitive operation.
Follow-up: What is the strongest control for an unauthorized refund? A payment executor that rejects it regardless of model output.
Q3: Design a guardrail system for a customer-service chatbot.
Sample answer: I would place controls around the entire workflow: authenticated and bounded input, authorized evidence retrieval, structured proposal validation, business and permission checks before actions, and result verification before the final message. Content and privacy checks cover the user-facing output, with buffering where exposure before validation is unacceptable. Each failure has an explicit fallback or handoff. I track false blocks, missed violations, latency, review load, and resolution quality. The design must keep legitimate support usable while protecting the consequences that matter most.
Follow-up: What happens if the safety service times out? Follow a defined policy for that specific check and capability, rather than universally skipping or blocking everything.
Q4: Why does valid JSON not make an action safe?
Sample answer: A schema can verify that an amount is numeric and a recipient is a string, but it does not prove that the amount is allowed or the recipient is authorized. I validate business rules and resource ownership using trusted server state, not only model-supplied fields. For a refund, I check order ownership, remaining refundable amount, current status, and required approval. I then make the operation safe to retry. Structure, authorization, and execution correctness are separate contracts.
Follow-up: Where should tenant identity come from? The authenticated request context, not a free-form model argument.
Q5: How do you know your guardrails improve the product?
Sample answer: I compare the system with and without the proposed control on representative and adversarial cases, measuring both prevented failures and legitimate work blocked. I assess reviewer load, added latency, user outcomes, and residual severe risks. I inspect disagreement cases and set a policy for appeals or correction. A rising block rate could indicate attacks, a traffic change, or a broken detector; it is not automatically evidence of better safety. I assign ownership for ongoing calibration and incident learning.
Follow-up: Should every guardrail share one threshold? No; their error consequences and measurement methods differ.
60-second interview answer
I use guardrails to enforce product constraints at input, retrieval, tool execution, and output. Deterministic rules handle permissions, limits, schemas, and business invariants; classifiers or model checks handle ambiguous content with measured error rates. I define what happens when a check fails or is unavailable, and I evaluate both missed harms and unnecessary blocks. Guardrails complement a good model and system design. They do not guarantee truth or safety, and they must not become the only barrier around consequential actions.