Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

RAG Fundamentals

By Anup Rai9 min readReviewed September 2026

Numerical examples are illustrative unless explicitly sourced.

Remember: Retrieve evidence, then answer with it.

Definition and purpose

Retrieval-augmented generation (RAG) combines retrieval of external information with generation conditioned on that information. In a question-answering application, the system finds relevant evidence at request time and supplies it to the answering model. A vector database is one implementation choice, not part of the definition.

The model's training does not automatically include an organization's current policies. Retrieval can supply them without updating model weights, provided the source collection is current and the correct evidence is found.

Suppose the question is “Can I return a refurbished laptop after 20 days?” The system needs the refurbished-product rule, any regional exception, and the applicable effective date. A general 30-day policy may be topically similar but insufficient. This is why the job is to find answering evidence, not merely text that resembles the question.

Build the evidence path before the answering path

First ingest the policy. Preserve its title, headings, tables, source ID, version, effective date, and access rules. Split it into searchable units called chunks. A chunk should retain the information needed to interpret its claim. If “except refurbished items” is split away from the return window, retrieval can produce a misleading fragment.

An embedding maps text into a numeric representation that supports learned similarity search. A query embedding can help find passages using different wording. Keyword search helps with exact product names, identifiers, and specialized terms. Hybrid search combines these signals. Since their raw scores may have different scales, use a defined combination such as rank fusion rather than adding arbitrary scores.

A reranker examines the query and candidate passages more closely to improve their ordering. It cannot recover evidence that was never retrieved unless the pipeline performs another search. It also adds processing cost and latency. Start with a measurable baseline and add it when the failure pattern warrants it.

At answer time, authenticate the user, search only the evidence they may access, select enough context to answer, and pass it to the model with the task and citation requirements. Verify that important claims match their cited evidence. When a required fact is missing, ask for clarification or abstain. An answer that cannot be supported should not become more confident merely because the interface expects fluent text.

Understand the variants without treating them as a ladder

A simple retrieve-then-generate pipeline can be sufficient for a bounded problem. An advanced pipeline may rewrite queries, combine searches, and rerank results. An agentic retrieval system lets a model choose searches or repeat retrieval based on what it finds. This can handle variable multi-step questions, but adds decisions, calls, and failure modes. A fixed pipeline can also use model components, so the distinction is control flow rather than “deterministic versus intelligent.”

Graph-based retrieval represents entities and relationships or builds summaries over groups of records. It can help some relationship and corpus-wide questions, but graph construction, provenance, and evaluation are additional work. No variant automatically solves missing, stale, or unauthorized data.

Fine-tuning changes model weights and can teach behavior or task patterns. RAG changes the evidence available at request time. Long context changes how much can be supplied in one call. These approaches can be combined; none removes the need to test whether the system answers the actual question correctly.

State the requirements

Functional requirements for a policy-answering service:

  1. Ingest and update the permitted policy collection.
  2. Retrieve evidence applicable to the question, region and effective date.
  3. Answer with inspectable source references.
  4. Ask for missing inputs or abstain when the evidence is insufficient.
  5. Apply deletions and permission changes to every derived retrieval path.

Non-functional requirements:

  1. Correctness and evidence completeness on representative questions.
  2. Tenant isolation and current authorization before evidence exposure.
  3. Defined freshness and deletion-propagation bounds.
  4. Complete-request latency, availability and cost per successful answer.
  5. Observable failures across ingestion, retrieval and generation.

The model usually receives retrieved text or other supported content, rather than raw embedding vectors. An embedding is a search representation, not the answer itself.

Two paths to draw

Architecture / visual model
flowchart TD D[Sources with versions and permissions] --> P[Parse and preserve structure] P --> C[Create searchable chunks and metadata] C --> I[Keyword and or vector index] Q[Authenticated question] --> R[Search permitted evidence] I --> R R --> K[Optionally rerank and pack context] K --> G[Generate answer with sources] G --> V[Validate or abstain]
Read diagram source
flowchart TD
    D[Sources with versions and permissions] --> P[Parse and preserve structure]
    P --> C[Create searchable chunks and metadata]
    C --> I[Keyword and or vector index]
    Q[Authenticated question] --> R[Search permitted evidence]
    I --> R
    R --> K[Optionally rerank and pack context]
    K --> G[Generate answer with sources]
    G --> V[Validate or abstain]

The ingestion path reads and updates documents. It also handles deletion and permission changes. The query path answers a user's current question under current access rules. Both need monitoring and ownership.

Build each stage for a reason

Stage Design question Failure to test
Parsing Preserve headings, lists, tables, and page references? OCR changes a decimal or loses a table header
Chunking What unit contains enough evidence to answer? Answer and qualification are split apart
Indexing Keyword, dense, or hybrid search? Exact SKU missed by semantic search
Retrieval How much evidence and which filters? Forbidden or obsolete document included
Reranking Does pairwise query-passage scoring improve ordering? Adds latency without useful gain
Context packing Which sources fit the budget? Duplicate passages crowd out conflicting evidence
Generation What answer and citation contract? A citation exists but does not support the claim

Chunk length and overlap are tuning parameters, not fixed best practices. Evaluate using your actual documents, questions, and model. Hybrid search combines lexical and semantic signals; score scales may differ, so use a defined fusion/ranking method.

Diagnose the policy example

Observation Likely problem Repair and tradeoff
Refurbished policy absent from the corpus Source coverage Add the authoritative source and its lifecycle owner
Rule indexed without its exception Parsing or chunk boundary Preserve structure; larger units may add noise
General policy ranks above the applicable rule Retrieval or ranking Evaluate lexical/dense signals and applicability filters
Correct rule retrieved, wrong answer generated Context use or generation Check claim support and compare evidence presentation
Region is unknown Missing user input Ask a targeted clarification rather than infer eligibility
Two effective versions disagree Source authority or freshness Reconcile versions or surface the unresolved conflict

The top result is not necessarily sufficient evidence. Before adding a component, locate the failing stage and define how its repair will be measured.

RAG, long context, and fine-tuning

Need Starting option Important limit
Small supplied document set Direct context Must fit useful context, cost, and permissions
Large or changing knowledge base Retrieval Ingestion, freshness, and relevance become dependencies
Repeated style, format, or task behavior Prompting; possibly fine-tuning Fine-tuning is not a live knowledge database
Structured exact facts or actions Database/API tool Authorization and business validation still required

These options can be combined. RAG is not a mandatory stage before fine-tuning, and a million-token window does not remove access, freshness, or relevance requirements. Choose from the failure you need to fix.

Operations and economics

Use source IDs, content versions, ACL metadata, ingestion status, and lineage. Propagate deletions to chunks, indexes, caches, and stored answers as required. Budget initial backfill and incremental updates; query cost alone understates ownership cost.

Cache only within an appropriate permission and freshness scope. Exact or semantic similarity does not prove that a cached answer is safe for another user. For high-risk updates, revalidate against authoritative state before exposure.

Interview checks

“Does RAG eliminate hallucinations?” No. It provides evidence; retrieval and generation can both fail.

“Why not retrieve more?” More context can add noise, contradictions, cost, and latency. Measure evidence coverage and final quality.

“What should we launch first?” A measured, permission-aware baseline with supported answers and abstention, then improve the largest observed failure category.

The original RAG paper establishes the retrieval-plus-generation idea; production permission and lifecycle requirements follow from the application. See RAG evaluation for debugging.

Follow the rank-fusion calculation

Hybrid retrieval often combines lists with different score scales. Reciprocal rank fusion uses RRF(d) = Σ 1 / (k + rank_i(d)) over lists containing document d. For k = 60, a document ranked first in one list and third in another scores 1/61 + 1/63 ≈ 0.03227. A document appearing only first in one list scores 1/61 ≈ 0.01639. Here k is the smoothing constant, not the number of returned results.

RRF rewards agreement without treating cosine similarity and BM25 as the same unit. It still requires candidate retrieval, permission enforcement, and relevance evaluation. Work through the reciprocal rank fusion derivation before tuning the result cutoff.

Interview questions with developed answers

Q1: Why use RAG when models have very large context windows?

Sample answer: A large window does not decide which data is relevant, current, or permitted for this user. Retrieval helps select evidence from a large or changing corpus and can reduce the material processed for each request. For a small bounded document set, direct context may be simpler and should be compared. I evaluate answer quality, access control, freshness, latency, and full cost rather than use a universal token threshold. Even when all documents fit, the model may fail to use a needed qualification among distracting material.

Follow-up: Does caching eliminate this tradeoff? It changes some processing and billing costs but does not guarantee relevance, permissions, or useful recall.

Q2: How does agentic RAG differ from an advanced retrieval pipeline?

Sample answer: An advanced pipeline follows a designed sequence such as rewrite, hybrid search, rerank, and generate. An agentic approach gives the model discretion over which source to search or whether to retrieve again. That flexibility can help multi-step questions with unpredictable information needs. It also requires budgets, tool permissions, loop handling, and evaluation of the whole trajectory. I would use it when measured failures show the fixed flow is insufficient, rather than assume more agent decisions always improve retrieval.

Follow-up: Can a simple RAG pipeline be production-ready? Yes, if it meets the workload's quality, security, and operating requirements.

Q3: How do you choose chunks and retrieval depth?

Sample answer: I begin with document structure and the evidence needed for real questions. Chunks should preserve headings, table relationships, and qualifications, while remaining selective enough for useful search. I test chunking and top-k together with retrieval coverage and final answer quality. More passages can improve coverage but also introduce noise, contradictions, and cost. I inspect actual failures and compare alternatives rather than memorize one chunk size. Updates and citations require stable source references alongside the text.

Follow-up: What does overlap help with? Boundary loss, at the cost of duplication; it does not repair a parser that destroyed a table.

Q4: Does a citation make a RAG answer trustworthy?

Sample answer: A citation is a reference, not proof. I check that it exists, that the cited passage supports the specific claim, and that the source is authoritative, current, and applicable to the user. The model can cite the wrong section or accurately quote an obsolete policy. I also check completeness: a cited return window without its exception can mislead. The product should expose evidence in a form the user or reviewer can inspect and abstain when the required support is missing.

Follow-up: What if two sources disagree? Apply an explicit authority and version policy or surface the unresolved conflict.

Q5: What does a complete first RAG release include?

Sample answer: A bounded corpus with an owner, reliable parsing and updates, permission-aware retrieval, evidence-based answers, and a clear abstention or handoff path. I include tests for common questions, missing evidence, stale policies, and unauthorized access, then monitor component failures and user outcomes. I track deletion propagation and cost as well as query latency. Starting simply means limiting scope and components while keeping these responsibilities explicit; it does not mean omitting operations until after launch.

Follow-up: Which enhancement comes next? The one that addresses the largest measured failure category.

60-second interview answer

Retrieval-augmented generation supplies a model with relevant external information at request time. I build an ingestion path that preserves source structure, versions, and permissions, and a query path that retrieves authorized evidence before generating an answer. I start with a simple baseline and add hybrid search or reranking only when failure analysis supports it. RAG helps with private or changing knowledge, but does not guarantee factual answers. I measure evidence retrieval, answer correctness, citations, abstention, freshness, latency, and cost separately.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Prompt Injection and Defense
NEXT LESSONChunking Strategies →

Explore the diagram