Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Case Study: Enterprise Knowledge Assistant

By Anup Rai14 min readReviewed September 2026

This is a hypothetical interview scenario. Traffic, latency, storage, staffing and costs are planning assumptions, not measured results. Connector behavior and model pricing have primary-source links.

Interview focus: retrieve evidence employees are allowed to use, distinguish current guidance from historical experience, and preserve that boundary through citations, caches and source changes.

60-second interview answer

I would start with permission-aware search over a small set of approved sources, then add evidence-backed answer synthesis. Each document would retain its source identity, access rules, version, owner and approval status. Retrieval would filter by the employee's access and revalidate selected evidence before model use and answer delivery. Current-policy questions would prefer the applicable approved edition; historical questions would keep their date context. Unsupported claims would become explicit gaps rather than confident prose. I would independently monitor connector freshness, access failures, citation support and time saved after verification, and give source owners a workflow to resolve stale or conflicting guidance.

Remember: Authorize → Retrieve → Check authority → Cite evidence → Revalidate and maintain.

Interview problem and scope

A consulting firm with 10,000 employees has two million documents across 15 systems, including SharePoint, Confluence and file shares. Consultants want to ask questions such as “How did earlier automotive projects reduce inventory?” or “Which methodology should this project use today?”

Those questions have different evidence requirements. Earlier project reports can describe historical experience without being approved current policy. A current methodology answer needs the applicable approved source, not simply the latest modified page. The assistant must not expose client-confidential work merely because the employee has the same seniority as its author.

Functional requirements

  1. Connect approved repositories and preserve source identity, version, ownership, access rules and lifecycle status.
  2. Search across permitted evidence, then synthesize answers with citations for substantive source-dependent claims.
  3. Distinguish current guidance, historical questions and unresolved conflicts or missing evidence.
  4. Reflect content edits, deletions, permission changes and group-membership changes.
  5. Let users open cited evidence through authorized links and report incorrect, stale or unsupported answers.
  6. Route knowledge-maintenance tasks to authorized owners without granting the assistant permission to publish policy.

Nonfunctional requirements

  1. Access: authorize evidence before it reaches the model, and recheck dependencies before returning an answer or cached answer.
  2. Freshness: target 95% of ordinary supported-source edits searchable within five minutes, with connector lag visible. Access revocation follows a separate authorization contract.
  3. Latency: target p95 under four seconds for a supported answer at the agreed burst load; return a bounded partial result or unavailable-source status when appropriate.
  4. Quality: measure retrieval coverage, claim support, correct policy edition and useful abstention separately.
  5. Audit: retain source versions, retrieval decisions and answer dependencies under the approved privacy/retention policy.
  6. Recovery: support replay, connector reconciliation and index replacement without broadening permissions.

Interview tip: “A private vector database” and “one namespace per company” do not enforce project, client or document permissions among employees of the same company.

Sizing assumptions

Assume 2,000 source tokens per document, six indexed passages per document and 50,000 answers per month.

Quantity Calculation Consequence
Source tokens 2M × 2,000 = 4B Initial parsing and embedding are batch work
Embedding tokens with 20% overlap allowance 4B × 1.2 = 4.8B Include overlap and repeated headers in measured usage
Passage vectors 2M × 6 = 12M Measure indexes and permission metadata as well as vectors
Raw vectors 12M × 1,536 dimensions × 4 bytes = 73.728 GB Two copies need 147.456 GB before ANN structures, text and backups
Monthly query rate 50,000 / 10,000 = 5 per employee on average This workload cannot support claims of hours saved by every employee each week
Burst assumption 20 queries/s at three seconds mean About 60 concurrent requests before headroom

If 1% of documents change daily, 20,000 documents × 2,000 tokens × 1.2 is 48M embedding tokens/day, or 1.44B per 30-day month. Permission-only changes should update access metadata and invalidate dependent caches without needlessly embedding unchanged text.

The baseline has three approved repositories, exact-keyword/full-text search, source permission enforcement and links to evidence. It already reduces fragmented searching while exposing connector and access-model problems.

Add hybrid retrieval and a reranker when evaluated queries require semantic matching. Add answer generation only after retrieval can find the right authorized passages. A polished generated answer cannot repair an index that omitted the relevant source or leaked restricted text.

Detailed architecture

Architecture / visual model
flowchart TB SOURCES[Approved repositories] --> CONN[Source-specific connectors and checkpoints] CONN --> JOB[(Durable ingest and access-change jobs)] JOB --> PARSE[Parse content and resolve source access metadata] PARSE --> CATALOG[(Versioned source and policy catalog)] PARSE --> INDEX[(Hybrid passage index)] USER[Employee query] --> ID[Authenticate and resolve current identity] ID --> RET[Permission-filtered retrieval] INDEX --> RET CATALOG --> RET RET --> CHECK[Authoritative access and version checks] CHECK --> RANK[Relevance and applicable authority] RANK --> ANSWER[Evidence-backed draft with claim citations] ANSWER --> VERIFY[Support checks and dependency revalidation] VERIFY --> OUT[Answer evidence gaps and conflicts] VERIFY -->|Store validated result| CACHE[(Scoped cache with source dependencies)] CACHE -->|Cache hit| VERIFY
Read diagram source
flowchart TB
    SOURCES[Approved repositories] --> CONN[Source-specific connectors and checkpoints]
    CONN --> JOB[(Durable ingest and access-change jobs)]
    JOB --> PARSE[Parse content and resolve source access metadata]
    PARSE --> CATALOG[(Versioned source and policy catalog)]
    PARSE --> INDEX[(Hybrid passage index)]
    USER[Employee query] --> ID[Authenticate and resolve current identity]
    ID --> RET[Permission-filtered retrieval]
    INDEX --> RET
    CATALOG --> RET
    RET --> CHECK[Authoritative access and version checks]
    CHECK --> RANK[Relevance and applicable authority]
    RANK --> ANSWER[Evidence-backed draft with claim citations]
    ANSWER --> VERIFY[Support checks and dependency revalidation]
    VERIFY --> OUT[Answer evidence gaps and conflicts]
    VERIFY -->|Store validated result| CACHE[(Scoped cache with source dependencies)]
    CACHE -->|Cache hit| VERIFY

Every cached response returns through revalidation. Titles, snippets, counts, suggested queries and source links are also disclosures: do not expose them through an unfiltered search, debug panel or error message.

APIs and records

Operation Contract
POST /answers Authenticated user, query, current/historical intent and optional requested date; client cannot choose arbitrary access groups
GET /answers/{id} Revalidate access to the answer and its source dependencies
GET /evidence/{id} Authorize the employee against the cited source/version; render only permitted material
POST /feedback Capture a specific answer/claim and feedback reason, with controlled visibility
Connector event ingestion Verify source, deduplicate event IDs and preserve revision/checkpoint ordering
Record Fields and purpose
Source object Repository and stable object ID, canonical link, owner, version, deletion state and sync health
Access record Source permission semantics, resolved principals, ACL/group versions and last authoritative check
Passage Source version, section/location, text hash, parser/chunker/embedding versions and access boundary
Guidance edition Approval authority, effective interval, jurisdiction/practice applicability and supersedes relation
Answer User scope, cited claims/passages, model/prompt versions, source dependencies and validation result
Connector checkpoint Source scope, continuation/cursor, durable progress, last success, retry state and reconciliation generation

Keep stable source IDs across renames and moves. A path or title is a useful display value, not a durable identity. Different repositories containing copied text can still have different access rights and authority; deduplication must not erase those distinctions.

Authorization is an active part of retrieval

An access-control list (ACL) describes allowed or denied principals under the source's rules. A classification label such as “confidential” does not by itself identify who may read a document. Seniority is not a substitute for client-team membership, ethical walls, legal restrictions or specific grants.

  1. Authenticate the employee and resolve current identity/group information through the trusted identity service.
  2. Restrict search candidates using the source-specific access model; do not let the LLM choose permission filters.
  3. Revalidate selected source objects before any passage enters model context. A broad connector service account's ability to read a file is not evidence that the employee can read it.
  4. Before returning an answer, recheck its dependencies. If a dependency is revoked or changed, discard or regenerate the affected answer; do not simply remove its citation while retaining its content.
  5. Apply the same checks to caches, exports, conversation history and evidence endpoints.

Choose an authoritative check that genuinely represents the employee: a source operation under an appropriate delegated identity, or an access service implementing the complete source semantics. A generic list of sharing grants may omit contextual behavior or inherited rules; confirm API limitations and required privileges. Microsoft Graph file permissions.

No independent copied index can promise instantaneous revocation across every source without an enforceable freshness/authorization contract. Define the check point and consistency guarantee for each connector. For strict sources, withhold evidence when current authorization cannot be established. If a source supports only a bounded-lag replicated policy, disclose that bound and decide whether the repository is eligible for synthesis at all. A five-minute content sync is not an access guarantee.

Cached answers retain their dependencies

The following application-level function checks a validated dependency manifest. For current-guidance answers, the access adapter verifies both current employee access and that the cited version is still the active usable source edition. Historical-answer caching needs a separate time-scoped edition check. It is not a vendor SDK call. A dependency check exception cannot fall through to returning the cached answer.

def dependencies_still_usable(user, dependencies, access):
    if not dependencies:
        return False
    try:
        return all(
            access.can_use_current_version(
                user, dep["source_id"], dep["source_version"]
            )
            for dep in dependencies
        )
    except (TimeoutError, ConnectionError):
        return False

Bind cache identity to tenant, employee or an explicitly equivalent access scope, query/intent, source dependencies and generation configuration. A cache key alone does not revoke previously cached data. Permission events invalidate known dependencies; authoritative checks cover missed invalidations. Content already legitimately viewed cannot be erased from a person's memory, but future application access and model reuse must follow the updated policy.

For mixed permissions, split passages only where the source or an approved policy defines an enforceable section boundary. Otherwise retain the whole document's restriction. Never merge a restricted passage with a broader one and assign the more permissive ACL.

Current policy and historical evidence need different ranking

Question Preferred source Correct treatment of older material
“What method must this team use now?” Applicable approved edition in effect today Exclude superseded guidance as current instruction
“What did the 2021 engagement do?” Authorized historical project evidence Retain the historical version and date context
“Why did the policy change?” Authorized old/new editions and recorded rationale Compare versions without presenting the old one as current
“What are teams experimenting with?” Clearly labeled permitted drafts and experiments Do not imply official endorsement

An old approved policy can remain valid. Yesterday's cosmetic edit does not make a draft authoritative. Metadata on applicability, effective dates and supersession belongs in the content-owner workflow; a model should not invent it from writing style.

This example scores current guidance only after access filtering and owner-maintained applicability resolution. semantic_score is a finite normalized relevance score from the chosen evaluation pipeline; it is not a probability of truth.

def current_guidance_score(doc, semantic_score, today):
    if (doc.status != "approved" or not doc.applies_to_question
            or doc.effective_date > today or doc.superseded):
        return None
    valid_until = getattr(doc, "valid_until", None)
    if valid_until is not None and today >= valid_until:
        return None
    if (isinstance(semantic_score, bool)
            or not isinstance(semantic_score, (int, float))
            or not 0 <= semantic_score <= 1):
        raise ValueError("Expected a normalized finite relevance score")
    return semantic_score

The optional validity interval is half-open: effective on its start date and no longer effective on valid_until. Historical queries use the requested time and corresponding edition catalog rather than today's superseded flag. Within equally applicable sources, evaluated recency or review-date signals can break ties. A recency half-life is a ranking preference, never an expiry rule for policy.

Source changes and index publication

Each connector needs its own delivery, permission and recovery contract.

Source Change mechanism Gaps to cover
SharePoint/OneDrive via Microsoft Graph Paged delta responses with saved continuation and delta links Sharing changes, deletions, parent effects and invalidated checkpoints
Confluence Cloud Supported Forge events plus incremental polling/reconciliation Page/space permission changes, lost events, app access and API limits
Managed file shares Change notifications/journals where available plus inventory scans Renames, ACL-only edits, missed events and disconnected hosts
Other repositories Source-specific event or polling API Do not assume one common delta-token protocol

For Microsoft Graph, follow the returned nextLink pages and save the completed deltaLink only after changes are durably represented. Handle deletions and the documented resynchronization cases. Sharing-change options help identify affected objects, but do not substitute for complete effective-access evaluation. Graph drive-item delta.

For a new Confluence Cloud app, use the current Forge event model rather than assuming legacy Connect webhooks are the default. Its documented events include page permission updates; an event's delivery and the app's content access are separate concerns. Track page and broader permission changes as supported, then reconcile. Forge Confluence events, Atlassian event access.

Architecture / visual model
flowchart LR EVENT[Content permission or deletion event] --> INBOX[(Deduplicated durable inbox)] INBOX --> CLASS{Change type} CLASS -->|Content| BUILD[Build new passages and metadata] BUILD --> READY[Verify searchable complete version] READY --> PUBLISH[Atomically switch active catalog edition] CLASS -->|Access or deletion| DENY[Update access or tombstone first] DENY --> INVALIDATE[Invalidate dependent answers and indexes] PUBLISH --> INVALIDATE SCAN[Periodic source reconciliation] --> INBOX
Read diagram source
flowchart LR
    EVENT[Content permission or deletion event] --> INBOX[(Deduplicated durable inbox)]
    INBOX --> CLASS{Change type}
    CLASS -->|Content| BUILD[Build new passages and metadata]
    BUILD --> READY[Verify searchable complete version]
    READY --> PUBLISH[Atomically switch active catalog edition]
    CLASS -->|Access or deletion| DENY[Update access or tombstone first]
    DENY --> INVALIDATE[Invalidate dependent answers and indexes]
    PUBLISH --> INVALIDATE
    SCAN[Periodic source reconciliation] --> INBOX

Separate materialization from publication. New passages become active only when parsing, access metadata and search visibility are ready. A catalog check suppresses superseded chunks while background cleanup removes old index entries. Apply deletions/restrictions through an immediate deny/tombstone path rather than waiting for an expensive re-embedding batch.

A missing file in a failed partial scan is not proof of deletion. Reconcile against a successfully completed, scoped inventory generation. Treat content hashes and permission versions separately: unchanged bytes can acquire a new restriction. During a rescan or outage, report actual source coverage instead of implying the entire knowledge base is current.

Evidence, conflicts and knowledge gaps

A citation must identify a source passage supporting the associated claim, not merely a document that mentions similar terms. Evaluate both citation correctness (does the cited passage support the claim?) and citation completeness (do substantive claims needing evidence have support?). Similarity scores and LLM self-confidence cannot establish either.

Situation Response behavior
Relevant permitted evidence supports the answer Answer with claim-level citations and the appropriate date context
Evidence answers only part of the question Give the supported part and clearly identify what remains unknown
Applicable approved policy conflicts with a project note Explain their different authority without blending them into a new policy
Two apparently authoritative editions disagree Show permitted conflicting passages and request owner resolution
Source temporarily unavailable State that the answer may be incomplete within the accessible sources
No permitted supporting evidence Say evidence was not found in material available to the user; do not reveal hidden document names

Contradiction detection is itself fallible. Compare claims with their scope, dates and conditions; two different client recommendations may both be correct. Cite the conflict and its context. Do not always select the newer page or fabricate a compromise.

A maintenance task should route to an authorized source owner, with only the evidence they may access. User feedback is a useful signal, not an automatic policy update or training label. Keep approval and publication permissions outside the answer-generating model.

Evaluate and repair the system

Create an adjudicated set spanning roles, client projects, group changes, languages, historical questions, policy conflicts, missing answers and injected document instructions. Use source/role pairs in access tests, including cases where the connector account can read more than the employee.

Failure test Required repair or behavior
Employee loses a project group after an answer is cached Revalidate dependencies and withhold the cached answer
Document moves to a restricted parent without text changes Refresh effective access and suppress stale passages
New draft duplicates the title of an approved policy Keep authority separate from recency and title similarity
Connector misses an edit or deletion Reconciliation repairs the index and dependent caches
Parser loses a table exception Report extraction coverage; do not claim the policy has no exception
Source text asks the model to reveal other projects Treat it as untrusted evidence, never an authorization instruction
Two users share a semantically similar query Reuse an answer only after proving the second user's access to every dependency

Measure retrieval recall for authorized relevant sources, support and completeness of citations, correct-edition selection, abstention usefulness, source lag and unauthorized disclosure tests. Record failures by connector and policy class rather than averaging them away. Zero observed access failures is a release test result, not proof of universal correctness.

Costs and realistic value

At an illustrative embedding rate of $0.02 per million tokens, the initial 4.8B tokens cost $96 and the 1.44B monthly refresh tokens cost $28.80, before parsing and retries. A 1,536-dimensional text-embedding-3-small representation is the sizing example; different embeddings change both quality and storage. OpenAI embedding pricing.

For 50,000 monthly answers with 4,000 input and 500 output tokens on Claude Sonnet 5 at $2/$10 per million tokens, one generation pass costs $650/month. Extra retrieval rewrites, checks and retries must be counted separately. Claude pricing.

Component Monthly planning amount
Embedding refresh $28.80
One answer-generation pass $650
Search, passage storage and replicas allowance $2,000
Connectors and synchronization allowance $1,000
Monitoring, access service and backups allowance $800
Partial infrastructure total $4,478.80
40 source-owner review hours at $100/hour $4,000
Partial operating total $8,478.80

Do not infer that every employee saves two hours weekly from five average queries per month. If 60% of 50,000 answers genuinely save two net minutes after verification, that is 1,000 hours of potential monthly capacity, valued at $100,000 using the assumed $100/hour. It is not automatically revenue, cash savings or a headcount reduction. Measure adoption, repeated failed searches, correction effort and how saved time is used.

Choice Benefit Cost or limitation
Filter plus authoritative recheck Protects model context despite stale candidate metadata Source latency, rate limits and lower availability when access cannot be established
Current-policy catalog Prevents a recent draft replacing approved guidance Requires accountable owners and version metadata
Claim-level evidence Easier verification and correction More careful extraction, answer structure and validation
Shared answer cache May reduce repeated generation cost Difficult permission equivalence and dependency invalidation
Phased connector rollout Finds source-specific correctness gaps early Coverage expands more slowly than a bulk import

Interview follow-ups

1. Why isn't filtering by the employee's department enough?

Documents may have client, project, individual, inherited and explicit-denial rules. Preserve the source's effective permission semantics. A broad organizational label cannot safely approximate them.

2. Does a five-minute sync satisfy revocation?

No. It is a content freshness target. Define authoritative access checks or an explicitly accepted permission-lag bound, revalidate cached dependencies and withhold evidence when that contract cannot be met.

3. Should newer documents always rank higher?

No. Current questions first require applicable approved guidance. Historical questions need the relevant past version. Recency is at most an evaluated ranking signal within the appropriate set.

4. What if no evidence is found?

Return a bounded gap statement and a permitted next step. Do not fabricate an answer or disclose the existence of inaccessible documents. Distinguish missing evidence from a known connector outage without leaking confidential metadata.

5. Can the model answer first and add citations afterward?

It should generate against an explicit evidence set with claim-to-passage references, then validate support. Adding plausible-looking links to unsupported prose does not make the answer grounded.

6. How do you recover a lost connector checkpoint?

Use the source's supported resynchronization path, rebuild a scoped inventory and reconcile only after complete enumeration. Retain deny/tombstone protections and do not infer deletion from a failed partial scan.

7. How do you demonstrate business value?

Measure task completion and net time saved on representative searches, including verification and failed answers. Combine that with real adoption and maintenance costs; avoid multiplying an assumed weekly saving by every employee.

Closing notes

This system succeeds when employees can find permitted, applicable and verifiable evidence faster. Start with secure search, then earn synthesis through source coverage and evaluation. Close the interview with explicit decisions about revocation consistency, authority ownership, source-outage behavior and how usefulness will be measured.

Related: RAG fundamentals, Access control, RAG evaluation.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Case Study: Real-Time Payment Fraud Decisions
NEXT LESSONCase Study: Expense Operations with a Computer-Use Agent →

Explore the diagram