Interview problem: build a system that helps analysts prepare company research reports from filings, earnings material and licensed research. Every published report requires an authorized analyst to approve its exact version.
The system's main challenge is data lineage: each number and factual claim must be traceable to its source, reporting context and transformations. An ensemble of agreeing models does not replace that evidence.
This is a hypothetical Learnastra exercise, not investment advice or a claim of measured production performance. Workload and cost figures are planning assumptions. Publication policies in the scenario belong to the firm; applicable legal obligations depend on its activity and jurisdiction.
1. Clarify the report and publication boundary
Ask whether the product drafts internal notes or externally distributed research, which entities/markets it covers, which data licenses apply, whether users ask historical “what was known then?” questions, and who may approve publication. Separate reported facts, computed metrics, analyst interpretation and forecasts.
Functional requirements
- Ingest approved filings, earnings releases/transcripts and licensed research, preserving source identities and versions.
- Extract financial facts with entity, period, unit, currency, accounting basis and source location.
- Reconcile amendments, duplicate disclosures and conflicting values without silently picking a convenient number.
- Calculate ratios and changes using deterministic, versioned code.
- Draft a report whose material factual claims link to accepted facts or source passages.
- Flag missing evidence, unsupported claims, contradictory sources and assumptions for analyst review.
- Require version-specific analyst approval before publishing; support corrections and withdrawal.
- Reproduce a report from its immutable evidence snapshot, configuration and calculation records.
Out of scope: autonomous trading, personalized investment recommendations, publication based only on a model score, and unrestricted redistribution of licensed source material.
Nonfunctional requirements
- Target 100 ordinary reports/day, with bursts after earnings releases; define page count and claim count distributions.
- Aim for under 30 minutes from a ready evidence packet to an approved ordinary report, including waiting and active review. Complex cases may need longer and must show that status.
- Aim for under $50 per approved ordinary report including allocated review and operating costs.
- Evaluate at least 99.5% correctly extracted facts on an independently labeled representative set; also report uncertainty and severe-error slices.
- Block publication with unresolved critical numerical errors, missing evidence or invalid approval. This is a release rule, not a proof of zero error.
- Enforce source entitlements, tenant access, retention policy and separation of author/reviewer roles where required by the firm's policy.
- Bound model/tool fan-out, elapsed time and spend; make interrupted jobs restartable without losing lineage.
A ±0.1% tolerance is ambiguous. Ask whether it means relative percentage error or 0.1 percentage points on a rate. Dates, currency and period errors cannot be made acceptable merely by meeting a numeric tolerance.
2. Separate facts from interpretations
| Claim class | Example | Evidence/check needed |
|---|---|---|
| Reported fact | “Annual revenue was $1.2 billion” | Filing, entity, period, unit/scale and source version |
| Derived metric | “Revenue increased 20%” | Two compatible revenue facts and a recorded formula |
| Quotation | A CEO's exact statement | Exact source passage, speaker and timestamp/page |
| Interpretation | “Margins improved because of product mix” | Supporting evidence and qualification of causal uncertainty |
| Forecast | “Revenue could grow under scenario A” | Explicit assumptions, method and scenario label |
| Valuation ratio | “P/E is 15” | Price timestamp, EPS definition/period and calculation |
An analyst opinion must not be relabeled as a company-reported fact. An earnings-call statement is not interchangeable with audited financial statements. A disclaimer does not supply missing evidence.
A small context error can become a large financial error
Suppose a table states amounts “in millions” and reports revenue as 1,200. Store the scale as part of extraction and normalize deliberately: the amount is 1.2 billion currency units, not 1,200 units.
Likewise, a nine-month year-to-date value is not a third-quarter value. Subtracting the six-month value from the nine-month value can derive the third quarter only when entity, consolidation scope, basis, units and restatement treatment are compatible. Do not subtract cumulative per-share figures blindly: their weighted-average share denominators may differ.
3. Start with structured facts and one drafting pass
Read diagram source
flowchart LR
S[Approved source packet] --> X[Extract structured facts and passages]
X --> C[Reconcile context and calculate metrics]
C --> D[One bounded drafting call]
D --> V[Check claims against sources and calculations]
V --> H[Analyst reviews exact version]
H --> P[Publish only approved artifact]
Use structured data when available. SEC EDGAR APIs expose submission history and XBRL facts. The aggregate XBRL APIs cover standard-taxonomy, whole-entity facts, so they are not a complete substitute for custom tags, segment details or the original filing. Calendar frames also require attention to different fiscal start/end dates.
Parse HTML/XBRL directly where reliable; use layout-aware extraction or a document model for uncovered tables and scans. Retain the visual source region for review. A model helps interpret difficult structure, but a JSON schema cannot establish that the reported value is true.
Diagnose before adding an ensemble
| Baseline failure | Targeted improvement | Benefit | Cost or remaining limitation |
|---|---|---|---|
| Misread table header/unit | Layout-aware extraction and source-region review | Fixes the relevant evidence boundary | Parsing/review effort |
| Wrong ratio or rounding | Deterministic decimal calculation | Reproducible arithmetic | Correct inputs/definitions still required |
| Draft overlooks a disclosed risk | Independent risk-focused review | May detect omissions | Additional calls; reviewer can miss the same fact |
| Synthesis introduces a new number | Regenerate claim ledger and verify final draft | Checks what will actually be published | Extra validation and possible rework |
| Reviewer disagreement | Inspect source and calculation, then analyst decision | Resolves evidence rather than voting | Human queue and longer turnaround |
| Stale report after amended filing | Dependency tracking and approval invalidation policy | Prevents silent reuse of obsolete facts | Reprocessing and correction workflow |
4. Detailed architecture and data lineage
Read diagram source
flowchart TD
subgraph INPUT[Controlled evidence intake]
F[Filings and amendments] --> ING[Connector with entitlement and cutoff checks]
E[Earnings material] --> ING
R[Licensed research] --> ING
ING --> OBJ[(Immutable source objects and hashes)]
OBJ --> PAR[Structured parser or bounded document extraction]
PAR --> FACT[(Candidate facts with full context)]
end
FACT --> REC[Reconcile source, period, unit and basis]
REC -->|Unresolved| EX[Analyst evidence queue]
REC -->|Accepted| AF[(Accepted fact snapshot)]
AF --> CALC[Versioned deterministic calculations]
CALC --> LED[(Facts, derived values and dependency graph)]
LED --> DRAFT[Drafting workflow]
DRAFT --> CLAIM[Final-draft claim inventory]
CLAIM --> VER[Numeric, citation, quote and support checks]
OBJ --> VER
LED --> VER
VER -->|Gaps| EX
EX -->|Corrected evidence or text| DRAFT
VER -->|Ready for review| UI[Analyst workbench]
UI --> AP[(Approval for exact artifact and evidence hash)]
AP --> PUB[Publication gate and correction registry]
PUB --> OUT[Approved report]
A report worker can run asynchronously through a durable queue. PostgreSQL can hold typed records and state transitions; an immutable object store holds filings and report artifacts. A search index helps locate passages, but it is not the authoritative numerical ledger. Model endpoints sit behind configured data-policy and budget controls.
Pin model/prompt/parser/calculation versions. Use current supported SDKs and test their structured-output behavior; do not rely on a hard-coded product name to establish extraction quality. Self-hosting versus an approved API depends on confidentiality, licensing, latency, capacity and operating cost.
Essential records
| Record | Required context |
|---|---|
| Source | Entity ID, filing/accession or publisher ID, published/accepted time, retrieved time, source hash, license scope |
| Fact | Concept, value, scale, currency/unit, instant or duration, start/end, fiscal period, accounting basis, segment/dimensions, source locator |
| Reconciliation | Candidate facts, selected version, reason, unresolved conflicts and reviewer |
| Calculation | Formula version, input fact IDs, exact result, display rounding and units |
| Claim | Exact draft span, class, supporting fact/passage IDs, verification status and severity |
| Report | Content hash, evidence snapshot, as-of cutoff, workflow/configuration IDs |
| Approval | Approver identity/role, report and evidence hashes, time, policy version |
Store when the fact applies and when the system knew it separately. A backtest or historical report cannot use an amendment published after its information cutoff without disclosing that choice. The newest downloaded value is not always the correct historical value.
5. Verify numbers with context and code
Worked calculations
Assume comparable annual revenue of $1,000M and $1,200M. Growth is (1,200 − 1,000) / 1,000 × 100 = 20%.
If operating margin changes from 18% to 20%, the difference is 2 percentage points, or 200 basis points. The relative increase in the margin rate is (20 − 18) / 18 × 100 ≈ 11.11%. These describe different quantities.
For a simple trailing P/E example, a $90 share price divided by $6 trailing diluted EPS is 15. Match share basis and currency, label the price timestamp, and do not mix trailing EPS with a forward estimate. With zero/negative earnings, the usual positive P/E comparison is not meaningful; show the underlying facts and the chosen analytical convention.
from decimal import Decimal, InvalidOperation
def positive_baseline_growth(previous, current):
"""Inputs must already be compatible facts; result is percent, unrounded."""
try:
old, new = Decimal(str(previous)), Decimal(str(current))
except InvalidOperation as exc:
raise ValueError("Use finite numeric facts") from exc
if not old.is_finite() or not new.is_finite():
raise ValueError("Use finite numeric facts")
if old <= 0:
return None # Route zero/negative-baseline wording to analyst review.
return (new - old) / old * Decimal("100")
assert positive_baseline_growth("1000", "1200") == Decimal("20")
assert positive_baseline_growth("100", "80") == Decimal("-20")
assert positive_baseline_growth("0", "10") is None
assert positive_baseline_growth("-10", "5") is None
The arithmetic formula is mathematically defined for a negative nonzero denominator, but “growth” wording can be misleading; the example deliberately routes it to review. Currency conversion, restatement matching and accounting definitions belong in upstream validation, not inside this arithmetic helper. Round only at the specified presentation step and retain the unrounded calculation.
6. Optional ensemble: measure incremental error detection
An ensemble combines multiple outputs or judgments. Extra model passes can reveal omissions and contradictions, but their errors may be correlated through shared training, evidence, prompts or parser mistakes.
| Stage | Optional extra work | Acceptance rule |
|---|---|---|
| Ambiguous extraction | Up to five independent candidate readings | Reconcile against source; unanimity alone is insufficient |
| Analysis | Quantitative, business-narrative and risk-focused drafts plus synthesis | Verify the final synthesized claims, including newly introduced ones |
| Verification | Separate numerical/context/contradiction reviewers | Evidence decides; disagreement triggers investigation |
| Writing review | Panel checks clarity, completeness and qualifications | Cannot waive factual checks or analyst approval |
Independent review before optional discussion
Give reviewers the claim and source evidence independently on the first pass. If they see one another's answers immediately, agreement may reflect anchoring. Permit at most two configured rounds for unresolved findings, then hold for an analyst when time or spend is exhausted. The next round must inspect an explicit disagreement, not simply ask the panel to be more confident.
Read diagram source
sequenceDiagram
participant W as Verification workflow
participant N as Numerical checker
participant A as Reviewer A
participant B as Reviewer B
participant H as Analyst
par Deterministic checks
W->>N: Claim, fact IDs and formula
N-->>W: Result and mismatch details
and Independent support review
W->>A: Claim and source passages
A-->>W: Finding with citations
and Independent contradiction review
W->>B: Claim and source passages
B-->>W: Finding with citations
end
alt Required check fails or reviewers disagree
W->>W: Bounded evidence investigation or one further round
W->>H: Unresolved findings and complete draft
else Required checks pass
W->>H: Draft ready for mandatory signoff
end
“Supported,” “derived,” “interpretation,” “unsupported” and “contradicted” are useful labels, but not a vote tally that produces truth. An inference may be reasonable and still require qualified wording. The claim extractor itself can miss an assertion; review the complete draft as well as the extracted checklist.
Evaluate the value of the added reviewers
On a protected set, compare the baseline with the ensemble using the same source snapshots and report tasks. Count additional true errors caught, new false alarms, errors missed by all reviewers, cost and analyst time. A claimed 98% detection rate without its dataset, denominator and method should not appear as a production fact.
If fact correctness were independently 99.5%, a report with 100 facts would have probability 0.995^100 ≈ 60.6% of all facts being correct. Real errors are often correlated, so this is an illustration of why per-fact accuracy is not report-level accuracy—not a forecast of this system's reliability.
7. Quality gate and approval mechanics
Use separate gates; a writing score cannot compensate for a wrong number. A statistical extraction target belongs to dataset evaluation. The system cannot look up a report's unknown “true accuracy” at runtime.
def next_stage(*, claims, unresolved_critical, checks):
required = ("numbers", "citations", "quotes", "required_disclosures")
if not claims or unresolved_critical:
return "resolve_evidence_gaps"
if any(checks.get(name) is not True for name in required):
return "resolve_evidence_gaps"
return "analyst_signoff" # Never an automatic publication permission.
passed = dict(numbers=True, citations=True, quotes=True, required_disclosures=True)
assert next_stage(claims=["c1"], unresolved_critical=[], checks=passed) == "analyst_signoff"
assert next_stage(claims=["c1"], unresolved_critical=["c1"], checks=passed) == "resolve_evidence_gaps"
assert next_stage(claims=[], unresolved_critical=[], checks=passed) == "resolve_evidence_gaps"
Checks must be based on actual evidence, not defaulted to true. Noncritical uncertain analysis must be corrected, removed or explicitly qualified under the review policy. The workbench displays source passages, calculation inputs, differences from prior versions and outstanding findings alongside the report.
Persist the review item before notifying a reviewer; use a retryable notification outbox so delivery failure does not lose the task. Deduplicate submissions by report version.
Publication verifies approval of the exact content/evidence hashes and current publication authority. Any material edit invalidates that approval. A new amendment marks dependent reports for reassessment; policy decides whether publication must be held, corrected or withdrawn. Reusing an old approval after regeneration is unsafe even if the report ID stays the same.
8. Cost, latency and analyst capacity
Assume a baseline machine pipeline costs $4.80 per report, allocated data/infrastructure adds $10, and an analyst spends 20 active minutes at $100/hour.
| Component | Cost per ordinary report |
|---|---|
| Parsing, model calls, numerical checks and retries | $4.80 |
| Allocated licensed data and infrastructure | $10.00 |
| Analyst review: 20/60 × $100 | $33.33 |
| Total | $48.13 |
These are hypothetical allowances. Actual pricing depends on input/output and reasoning tokens, image processing, context tiers, caching, provider contracts and the number of passes. An ensemble costing $8 more would take the same report above the $50 target before any extra review time. It may still be justified by error reduction, but the tradeoff must be explicit.
If machine processing takes eight minutes and review takes 20, only two minutes remain for queueing under the 30-minute goal. At 100 reports/day, review alone needs 100 × 20 / 60 ≈ 33.3 active analyst hours/day. Five reviewers with six active review hours each are insufficient at that handling time; six provide 36 hours with little burst headroom.
A faster drafting model will not fix an overloaded analyst queue. Compare narrower report scope, better evidence presentation, prioritized review and staffing. Do not meet the target by skipping required approval.
9. Failure handling, evaluation and rollout
| Failure | Correct response |
|---|---|
| Filing API delayed/unavailable | Retry within source-access limits; use an explicitly identified snapshot or hold |
| Conflicting facts | Preserve both contexts; investigate amendment, period, units and dimensions |
| Parser drops a table header | Reject ambiguous extraction and show the source region |
| Model invents a source ID | Fail citation validation; repair within budget or hold |
| New claim appears in synthesis | Rebuild the claim inventory and rerun relevant checks |
| Reviewer cannot finish before deadline | Mark delayed; do not auto-publish |
| Approval exists for older text | Require review of the current artifact |
| Published report becomes materially wrong | Correction/withdrawal process with linked versions and notifications under firm policy |
Build evaluations around facts and reports: extraction accuracy, unit/period errors, citation support, calculation correctness, missed claims, analyst corrections and final report-level severe errors. Slice by document layout, currency, accounting basis, custom tags, amendments and languages. Include adversarial instructions embedded in documents; they remain untrusted source data.
Start with internal drafts for a narrow company/report class. Run source-checking and analyst review before introducing optional multi-model passes. Assign data stewardship, calculation ownership, model evaluation, review queue operations and publication control explicitly. Retention and disclosure rules must be provided by the firm's responsible policy/compliance owners, not inferred from an LLM score.
Interview questions and developed answers
1. Why not use five-model unanimity for every number? Shared extraction errors can make all five agree on the wrong row, period or unit. Use authoritative structured facts and deterministic calculations first. Extra readings are useful when they expose ambiguity, but source reconciliation decides acceptance.
2. How do you avoid look-ahead bias in historical analysis? Set an information cutoff and preserve both effective periods and publication/knowledge times. Select only evidence available under that cutoff. Later amendments can inform a separate corrected analysis but must not silently enter the historical one.
3. A claim cites the right filing. Is that enough? No. Check the exact passage and context, including units, fiscal period, segment and accounting basis. A citation to a long filing does not establish support for a particular assertion.
4. What would you do with a negative growth baseline? The arithmetic may be defined but ordinary growth language can be misleading. Present the underlying change and route the interpretation through the specified analyst convention. Do not hide the sign or silently divide by its absolute value.
5. When would you add a second reviewer model? When protected evaluation shows it detects meaningful errors beyond deterministic checks and the existing reviewer, at an acceptable false-positive, latency and labor cost. Distinct vendors alone do not guarantee independent failures.
6. How do you protect the approved report during publication? Publish the immutable artifact whose hash and evidence snapshot match stored approval, recheck publication authority, and invalidate approval after material changes. Track amendments and corrections as new versions.
Closing remarks and recall table
| Remember | What must remain attached |
|---|---|
| Number | Entity, period, unit, basis and source |
| Calculation | Formula version and input fact IDs |
| Claim | Evidence and its classification |
| Report | Immutable text and evidence snapshot |
| Approval | Exact version and authorized reviewer |
60-second interview answer
I would automate evidence preparation and drafting while keeping arithmetic and publication authority outside the model. Every number retains its entity, period, units and source; every derived metric records its inputs and formula. Reviewers inspect the exact report and evidence snapshot before publication. Additional models are useful only when evaluation shows incremental error detection. I would budget analyst time and queueing, preserve historical information cutoffs, and maintain an explicit correction path after publication.
Remember: Source → Context → Calculation → Claim → Exact-version approval.