Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Design a Source-Verified Financial Research Assistant

By Anup Rai12 min readReviewed September 2026

Interview problem: build a system that helps analysts prepare company research reports from filings, earnings material and licensed research. Every published report requires an authorized analyst to approve its exact version.

The system's main challenge is data lineage: each number and factual claim must be traceable to its source, reporting context and transformations. An ensemble of agreeing models does not replace that evidence.

This is a hypothetical Learnastra exercise, not investment advice or a claim of measured production performance. Workload and cost figures are planning assumptions. Publication policies in the scenario belong to the firm; applicable legal obligations depend on its activity and jurisdiction.

1. Clarify the report and publication boundary

Ask whether the product drafts internal notes or externally distributed research, which entities/markets it covers, which data licenses apply, whether users ask historical “what was known then?” questions, and who may approve publication. Separate reported facts, computed metrics, analyst interpretation and forecasts.

Functional requirements

  1. Ingest approved filings, earnings releases/transcripts and licensed research, preserving source identities and versions.
  2. Extract financial facts with entity, period, unit, currency, accounting basis and source location.
  3. Reconcile amendments, duplicate disclosures and conflicting values without silently picking a convenient number.
  4. Calculate ratios and changes using deterministic, versioned code.
  5. Draft a report whose material factual claims link to accepted facts or source passages.
  6. Flag missing evidence, unsupported claims, contradictory sources and assumptions for analyst review.
  7. Require version-specific analyst approval before publishing; support corrections and withdrawal.
  8. Reproduce a report from its immutable evidence snapshot, configuration and calculation records.

Out of scope: autonomous trading, personalized investment recommendations, publication based only on a model score, and unrestricted redistribution of licensed source material.

Nonfunctional requirements

  1. Target 100 ordinary reports/day, with bursts after earnings releases; define page count and claim count distributions.
  2. Aim for under 30 minutes from a ready evidence packet to an approved ordinary report, including waiting and active review. Complex cases may need longer and must show that status.
  3. Aim for under $50 per approved ordinary report including allocated review and operating costs.
  4. Evaluate at least 99.5% correctly extracted facts on an independently labeled representative set; also report uncertainty and severe-error slices.
  5. Block publication with unresolved critical numerical errors, missing evidence or invalid approval. This is a release rule, not a proof of zero error.
  6. Enforce source entitlements, tenant access, retention policy and separation of author/reviewer roles where required by the firm's policy.
  7. Bound model/tool fan-out, elapsed time and spend; make interrupted jobs restartable without losing lineage.

A ±0.1% tolerance is ambiguous. Ask whether it means relative percentage error or 0.1 percentage points on a rate. Dates, currency and period errors cannot be made acceptable merely by meeting a numeric tolerance.

2. Separate facts from interpretations

Claim class Example Evidence/check needed
Reported fact “Annual revenue was $1.2 billion” Filing, entity, period, unit/scale and source version
Derived metric “Revenue increased 20%” Two compatible revenue facts and a recorded formula
Quotation A CEO's exact statement Exact source passage, speaker and timestamp/page
Interpretation “Margins improved because of product mix” Supporting evidence and qualification of causal uncertainty
Forecast “Revenue could grow under scenario A” Explicit assumptions, method and scenario label
Valuation ratio “P/E is 15” Price timestamp, EPS definition/period and calculation

An analyst opinion must not be relabeled as a company-reported fact. An earnings-call statement is not interchangeable with audited financial statements. A disclaimer does not supply missing evidence.

A small context error can become a large financial error

Suppose a table states amounts “in millions” and reports revenue as 1,200. Store the scale as part of extraction and normalize deliberately: the amount is 1.2 billion currency units, not 1,200 units.

Likewise, a nine-month year-to-date value is not a third-quarter value. Subtracting the six-month value from the nine-month value can derive the third quarter only when entity, consolidation scope, basis, units and restatement treatment are compatible. Do not subtract cumulative per-share figures blindly: their weighted-average share denominators may differ.

3. Start with structured facts and one drafting pass

Architecture / visual model
flowchart LR S[Approved source packet] --> X[Extract structured facts and passages] X --> C[Reconcile context and calculate metrics] C --> D[One bounded drafting call] D --> V[Check claims against sources and calculations] V --> H[Analyst reviews exact version] H --> P[Publish only approved artifact]
Read diagram source
flowchart LR
    S[Approved source packet] --> X[Extract structured facts and passages]
    X --> C[Reconcile context and calculate metrics]
    C --> D[One bounded drafting call]
    D --> V[Check claims against sources and calculations]
    V --> H[Analyst reviews exact version]
    H --> P[Publish only approved artifact]

Use structured data when available. SEC EDGAR APIs expose submission history and XBRL facts. The aggregate XBRL APIs cover standard-taxonomy, whole-entity facts, so they are not a complete substitute for custom tags, segment details or the original filing. Calendar frames also require attention to different fiscal start/end dates.

Parse HTML/XBRL directly where reliable; use layout-aware extraction or a document model for uncovered tables and scans. Retain the visual source region for review. A model helps interpret difficult structure, but a JSON schema cannot establish that the reported value is true.

Diagnose before adding an ensemble

Baseline failure Targeted improvement Benefit Cost or remaining limitation
Misread table header/unit Layout-aware extraction and source-region review Fixes the relevant evidence boundary Parsing/review effort
Wrong ratio or rounding Deterministic decimal calculation Reproducible arithmetic Correct inputs/definitions still required
Draft overlooks a disclosed risk Independent risk-focused review May detect omissions Additional calls; reviewer can miss the same fact
Synthesis introduces a new number Regenerate claim ledger and verify final draft Checks what will actually be published Extra validation and possible rework
Reviewer disagreement Inspect source and calculation, then analyst decision Resolves evidence rather than voting Human queue and longer turnaround
Stale report after amended filing Dependency tracking and approval invalidation policy Prevents silent reuse of obsolete facts Reprocessing and correction workflow

4. Detailed architecture and data lineage

Architecture / visual model
flowchart TD subgraph INPUT[Controlled evidence intake] F[Filings and amendments] --> ING[Connector with entitlement and cutoff checks] E[Earnings material] --> ING R[Licensed research] --> ING ING --> OBJ[(Immutable source objects and hashes)] OBJ --> PAR[Structured parser or bounded document extraction] PAR --> FACT[(Candidate facts with full context)] end FACT --> REC[Reconcile source, period, unit and basis] REC -->|Unresolved| EX[Analyst evidence queue] REC -->|Accepted| AF[(Accepted fact snapshot)] AF --> CALC[Versioned deterministic calculations] CALC --> LED[(Facts, derived values and dependency graph)] LED --> DRAFT[Drafting workflow] DRAFT --> CLAIM[Final-draft claim inventory] CLAIM --> VER[Numeric, citation, quote and support checks] OBJ --> VER LED --> VER VER -->|Gaps| EX EX -->|Corrected evidence or text| DRAFT VER -->|Ready for review| UI[Analyst workbench] UI --> AP[(Approval for exact artifact and evidence hash)] AP --> PUB[Publication gate and correction registry] PUB --> OUT[Approved report]
Read diagram source
flowchart TD
    subgraph INPUT[Controlled evidence intake]
        F[Filings and amendments] --> ING[Connector with entitlement and cutoff checks]
        E[Earnings material] --> ING
        R[Licensed research] --> ING
        ING --> OBJ[(Immutable source objects and hashes)]
        OBJ --> PAR[Structured parser or bounded document extraction]
        PAR --> FACT[(Candidate facts with full context)]
    end
    FACT --> REC[Reconcile source, period, unit and basis]
    REC -->|Unresolved| EX[Analyst evidence queue]
    REC -->|Accepted| AF[(Accepted fact snapshot)]
    AF --> CALC[Versioned deterministic calculations]
    CALC --> LED[(Facts, derived values and dependency graph)]
    LED --> DRAFT[Drafting workflow]
    DRAFT --> CLAIM[Final-draft claim inventory]
    CLAIM --> VER[Numeric, citation, quote and support checks]
    OBJ --> VER
    LED --> VER
    VER -->|Gaps| EX
    EX -->|Corrected evidence or text| DRAFT
    VER -->|Ready for review| UI[Analyst workbench]
    UI --> AP[(Approval for exact artifact and evidence hash)]
    AP --> PUB[Publication gate and correction registry]
    PUB --> OUT[Approved report]

A report worker can run asynchronously through a durable queue. PostgreSQL can hold typed records and state transitions; an immutable object store holds filings and report artifacts. A search index helps locate passages, but it is not the authoritative numerical ledger. Model endpoints sit behind configured data-policy and budget controls.

Pin model/prompt/parser/calculation versions. Use current supported SDKs and test their structured-output behavior; do not rely on a hard-coded product name to establish extraction quality. Self-hosting versus an approved API depends on confidentiality, licensing, latency, capacity and operating cost.

Essential records

Record Required context
Source Entity ID, filing/accession or publisher ID, published/accepted time, retrieved time, source hash, license scope
Fact Concept, value, scale, currency/unit, instant or duration, start/end, fiscal period, accounting basis, segment/dimensions, source locator
Reconciliation Candidate facts, selected version, reason, unresolved conflicts and reviewer
Calculation Formula version, input fact IDs, exact result, display rounding and units
Claim Exact draft span, class, supporting fact/passage IDs, verification status and severity
Report Content hash, evidence snapshot, as-of cutoff, workflow/configuration IDs
Approval Approver identity/role, report and evidence hashes, time, policy version

Store when the fact applies and when the system knew it separately. A backtest or historical report cannot use an amendment published after its information cutoff without disclosing that choice. The newest downloaded value is not always the correct historical value.

5. Verify numbers with context and code

Worked calculations

Assume comparable annual revenue of $1,000M and $1,200M. Growth is (1,200 − 1,000) / 1,000 × 100 = 20%.

If operating margin changes from 18% to 20%, the difference is 2 percentage points, or 200 basis points. The relative increase in the margin rate is (20 − 18) / 18 × 100 ≈ 11.11%. These describe different quantities.

For a simple trailing P/E example, a $90 share price divided by $6 trailing diluted EPS is 15. Match share basis and currency, label the price timestamp, and do not mix trailing EPS with a forward estimate. With zero/negative earnings, the usual positive P/E comparison is not meaningful; show the underlying facts and the chosen analytical convention.

from decimal import Decimal, InvalidOperation

def positive_baseline_growth(previous, current):
    """Inputs must already be compatible facts; result is percent, unrounded."""
    try:
        old, new = Decimal(str(previous)), Decimal(str(current))
    except InvalidOperation as exc:
        raise ValueError("Use finite numeric facts") from exc
    if not old.is_finite() or not new.is_finite():
        raise ValueError("Use finite numeric facts")
    if old <= 0:
        return None  # Route zero/negative-baseline wording to analyst review.
    return (new - old) / old * Decimal("100")

assert positive_baseline_growth("1000", "1200") == Decimal("20")
assert positive_baseline_growth("100", "80") == Decimal("-20")
assert positive_baseline_growth("0", "10") is None
assert positive_baseline_growth("-10", "5") is None

The arithmetic formula is mathematically defined for a negative nonzero denominator, but “growth” wording can be misleading; the example deliberately routes it to review. Currency conversion, restatement matching and accounting definitions belong in upstream validation, not inside this arithmetic helper. Round only at the specified presentation step and retain the unrounded calculation.

6. Optional ensemble: measure incremental error detection

An ensemble combines multiple outputs or judgments. Extra model passes can reveal omissions and contradictions, but their errors may be correlated through shared training, evidence, prompts or parser mistakes.

Stage Optional extra work Acceptance rule
Ambiguous extraction Up to five independent candidate readings Reconcile against source; unanimity alone is insufficient
Analysis Quantitative, business-narrative and risk-focused drafts plus synthesis Verify the final synthesized claims, including newly introduced ones
Verification Separate numerical/context/contradiction reviewers Evidence decides; disagreement triggers investigation
Writing review Panel checks clarity, completeness and qualifications Cannot waive factual checks or analyst approval

Independent review before optional discussion

Give reviewers the claim and source evidence independently on the first pass. If they see one another's answers immediately, agreement may reflect anchoring. Permit at most two configured rounds for unresolved findings, then hold for an analyst when time or spend is exhausted. The next round must inspect an explicit disagreement, not simply ask the panel to be more confident.

Architecture / visual model
sequenceDiagram participant W as Verification workflow participant N as Numerical checker participant A as Reviewer A participant B as Reviewer B participant H as Analyst par Deterministic checks W->>N: Claim, fact IDs and formula N-->>W: Result and mismatch details and Independent support review W->>A: Claim and source passages A-->>W: Finding with citations and Independent contradiction review W->>B: Claim and source passages B-->>W: Finding with citations end alt Required check fails or reviewers disagree W->>W: Bounded evidence investigation or one further round W->>H: Unresolved findings and complete draft else Required checks pass W->>H: Draft ready for mandatory signoff end
Read diagram source
sequenceDiagram
    participant W as Verification workflow
    participant N as Numerical checker
    participant A as Reviewer A
    participant B as Reviewer B
    participant H as Analyst
    par Deterministic checks
        W->>N: Claim, fact IDs and formula
        N-->>W: Result and mismatch details
    and Independent support review
        W->>A: Claim and source passages
        A-->>W: Finding with citations
    and Independent contradiction review
        W->>B: Claim and source passages
        B-->>W: Finding with citations
    end
    alt Required check fails or reviewers disagree
        W->>W: Bounded evidence investigation or one further round
        W->>H: Unresolved findings and complete draft
    else Required checks pass
        W->>H: Draft ready for mandatory signoff
    end

“Supported,” “derived,” “interpretation,” “unsupported” and “contradicted” are useful labels, but not a vote tally that produces truth. An inference may be reasonable and still require qualified wording. The claim extractor itself can miss an assertion; review the complete draft as well as the extracted checklist.

Evaluate the value of the added reviewers

On a protected set, compare the baseline with the ensemble using the same source snapshots and report tasks. Count additional true errors caught, new false alarms, errors missed by all reviewers, cost and analyst time. A claimed 98% detection rate without its dataset, denominator and method should not appear as a production fact.

If fact correctness were independently 99.5%, a report with 100 facts would have probability 0.995^100 ≈ 60.6% of all facts being correct. Real errors are often correlated, so this is an illustration of why per-fact accuracy is not report-level accuracy—not a forecast of this system's reliability.

7. Quality gate and approval mechanics

Use separate gates; a writing score cannot compensate for a wrong number. A statistical extraction target belongs to dataset evaluation. The system cannot look up a report's unknown “true accuracy” at runtime.

def next_stage(*, claims, unresolved_critical, checks):
    required = ("numbers", "citations", "quotes", "required_disclosures")
    if not claims or unresolved_critical:
        return "resolve_evidence_gaps"
    if any(checks.get(name) is not True for name in required):
        return "resolve_evidence_gaps"
    return "analyst_signoff"  # Never an automatic publication permission.

passed = dict(numbers=True, citations=True, quotes=True, required_disclosures=True)
assert next_stage(claims=["c1"], unresolved_critical=[], checks=passed) == "analyst_signoff"
assert next_stage(claims=["c1"], unresolved_critical=["c1"], checks=passed) == "resolve_evidence_gaps"
assert next_stage(claims=[], unresolved_critical=[], checks=passed) == "resolve_evidence_gaps"

Checks must be based on actual evidence, not defaulted to true. Noncritical uncertain analysis must be corrected, removed or explicitly qualified under the review policy. The workbench displays source passages, calculation inputs, differences from prior versions and outstanding findings alongside the report.

Persist the review item before notifying a reviewer; use a retryable notification outbox so delivery failure does not lose the task. Deduplicate submissions by report version.

Publication verifies approval of the exact content/evidence hashes and current publication authority. Any material edit invalidates that approval. A new amendment marks dependent reports for reassessment; policy decides whether publication must be held, corrected or withdrawn. Reusing an old approval after regeneration is unsafe even if the report ID stays the same.

8. Cost, latency and analyst capacity

Assume a baseline machine pipeline costs $4.80 per report, allocated data/infrastructure adds $10, and an analyst spends 20 active minutes at $100/hour.

Component Cost per ordinary report
Parsing, model calls, numerical checks and retries $4.80
Allocated licensed data and infrastructure $10.00
Analyst review: 20/60 × $100 $33.33
Total $48.13

These are hypothetical allowances. Actual pricing depends on input/output and reasoning tokens, image processing, context tiers, caching, provider contracts and the number of passes. An ensemble costing $8 more would take the same report above the $50 target before any extra review time. It may still be justified by error reduction, but the tradeoff must be explicit.

If machine processing takes eight minutes and review takes 20, only two minutes remain for queueing under the 30-minute goal. At 100 reports/day, review alone needs 100 × 20 / 60 ≈ 33.3 active analyst hours/day. Five reviewers with six active review hours each are insufficient at that handling time; six provide 36 hours with little burst headroom.

A faster drafting model will not fix an overloaded analyst queue. Compare narrower report scope, better evidence presentation, prioritized review and staffing. Do not meet the target by skipping required approval.

9. Failure handling, evaluation and rollout

Failure Correct response
Filing API delayed/unavailable Retry within source-access limits; use an explicitly identified snapshot or hold
Conflicting facts Preserve both contexts; investigate amendment, period, units and dimensions
Parser drops a table header Reject ambiguous extraction and show the source region
Model invents a source ID Fail citation validation; repair within budget or hold
New claim appears in synthesis Rebuild the claim inventory and rerun relevant checks
Reviewer cannot finish before deadline Mark delayed; do not auto-publish
Approval exists for older text Require review of the current artifact
Published report becomes materially wrong Correction/withdrawal process with linked versions and notifications under firm policy

Build evaluations around facts and reports: extraction accuracy, unit/period errors, citation support, calculation correctness, missed claims, analyst corrections and final report-level severe errors. Slice by document layout, currency, accounting basis, custom tags, amendments and languages. Include adversarial instructions embedded in documents; they remain untrusted source data.

Start with internal drafts for a narrow company/report class. Run source-checking and analyst review before introducing optional multi-model passes. Assign data stewardship, calculation ownership, model evaluation, review queue operations and publication control explicitly. Retention and disclosure rules must be provided by the firm's responsible policy/compliance owners, not inferred from an LLM score.

Interview questions and developed answers

1. Why not use five-model unanimity for every number? Shared extraction errors can make all five agree on the wrong row, period or unit. Use authoritative structured facts and deterministic calculations first. Extra readings are useful when they expose ambiguity, but source reconciliation decides acceptance.

2. How do you avoid look-ahead bias in historical analysis? Set an information cutoff and preserve both effective periods and publication/knowledge times. Select only evidence available under that cutoff. Later amendments can inform a separate corrected analysis but must not silently enter the historical one.

3. A claim cites the right filing. Is that enough? No. Check the exact passage and context, including units, fiscal period, segment and accounting basis. A citation to a long filing does not establish support for a particular assertion.

4. What would you do with a negative growth baseline? The arithmetic may be defined but ordinary growth language can be misleading. Present the underlying change and route the interpretation through the specified analyst convention. Do not hide the sign or silently divide by its absolute value.

5. When would you add a second reviewer model? When protected evaluation shows it detects meaningful errors beyond deterministic checks and the existing reviewer, at an acceptable false-positive, latency and labor cost. Distinct vendors alone do not guarantee independent failures.

6. How do you protect the approved report during publication? Publish the immutable artifact whose hash and evidence snapshot match stored approval, recheck publication authority, and invalidate approval after material changes. Track amendments and corrections as new versions.

Closing remarks and recall table

Remember What must remain attached
Number Entity, period, unit, basis and source
Calculation Formula version and input fact IDs
Claim Evidence and its classification
Report Immutable text and evidence snapshot
Approval Exact version and authorized reviewer

60-second interview answer

I would automate evidence preparation and drafting while keeping arithmetic and publication authority outside the model. Every number retains its entity, period, units and source; every derived metric records its inputs and formula. Reviewers inspect the exact report and evidence snapshot before publication. Additional models are useful only when evaluation shows incremental error detection. I would budget analyst time and queueing, preserve historical information cutoffs, and maintain an explicit correction path after publication.

Remember: Source → Context → Calculation → Claim → Exact-version approval.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Design a Conversational Customer-Support Agent
NEXT LESSONDesign an IDE Code Assistant →

Explore the diagram