Interview problem: answer questions about company filings, licensed news and market commentary with timely evidence, explicit coverage and a short response deadline.
Freshness is the delay between a source version becoming available and that version becoming usable by the service. It is different from recency, which describes how old an event is. An hour-old article indexed immediately is fresh ingestion of an older event; a one-minute-old article stuck in a queue is not yet searchable.
This is a hypothetical Learnastra interview scenario. Workload, latency and infrastructure budgets are planning assumptions. The service supports research, not autonomous trading or a guarantee that every external source is true.
1. Requirements and scope
Ask which sources are licensed, which companies and languages matter, how corrections work, whether users need historical “what was known then” answers, and whether sentiment means a summary of examples or a measured distribution over a defined population.
Functional requirements
- Collect eligible updates, corrections and deletions from approved sources.
- Search by company, ticker, terms, meaning and a requested time window.
- Return a concise answer with claim-level references and evidence timestamps.
- Compute financial values from typed, attributed facts rather than guessing from prose.
- Show missing sources, stale indexes and insufficient evidence explicitly.
- Support a separate, slower investigation mode for questions that cannot fit the fast path.
- Enforce source entitlements and retain an auditable record of which versions supported an answer.
Nonfunctional requirements
- Support 10,000 connected users and an average 50,000 queries/hour; determine burstiness separately.
- Target p95 complete, checked answers below three seconds for the bounded fast-query class.
- Target p99 source-availability-to-searchability below five minutes for eligible updates, per source and index. Report exclusions and outages separately rather than hiding them from the denominator.
- Propose 99.9% monthly availability for the fast path, while defining degraded search-only responses separately from successful synthesized answers.
- Preserve version ordering, source rights and deletion policy during retries and partial failure.
- Bound source polling, retrieved evidence, generation tokens and per-query spend.
- Measure unsupported claims, numeric errors and evidence coverage; citations alone cannot prove correctness.
Tip: “Real time” needs a clock and a percentile. A five-minute ingestion objective may be satisfied by frequent incremental polling; it does not automatically require a continuously streaming API.
2. Establish scale before choosing components
At 50,000 queries/hour, average load is 50,000 / 3,600 ≈ 13.9 queries/second. A planning peak of 10× gives about 139/s. If mean residence at that peak is two seconds, Little's Law implies about 278 in-flight queries under stable conditions. The 10,000 connected users are not 10,000 simultaneous model calls.
For an independent ingestion example, assume 1M new or changed documents/day, averaging 8KB normalized text and ten passages each:
| Quantity | Calculation | Result |
|---|---|---|
| Normalized text/day | 1M × 8KB | 8GB |
| Passages/day | 1M × 10 | 10M |
| Raw 1,024-dimensional float32 vectors/day | 10M × 1,024 × 4 bytes | 40.96GB |
| Thirty-day raw vector retention | 40.96GB × 30 | 1.2288TB |
| Two total copies | 1.2288TB × 2 | 2.4576TB |
Decimal units are used. These figures exclude index structures, payloads, source objects and superseded-version cleanup. Document changes can reuse unchanged passage embeddings if the exact text and embedding configuration match. A short “recent” index improves query locality; it does not eliminate archival or correction handling.
3. Begin with a searchable source catalog
The baseline is an incremental collector, a durable source-version store, a keyword index and a query API. Return ranked passages with timestamps first. Add answer generation only after retrieval and citation behavior are measurable.
For a small source set, a database-backed job queue and scheduled polling can be enough. A durable event log becomes useful when replay, several independent consumers, larger bursts or longer retention justify its operational cost. See pipeline and serving design.
| Baseline weakness | Change | Benefit | Cost or new risk |
|---|---|---|---|
| Exact terms miss paraphrases | Add semantic retrieval and rank fusion | Better recall for meaning-based questions | Embedding cost, second-index lag |
| One source correction leaves stale passages | Versioned publication and candidate validation | Prevents mixed revisions | Extra state and readiness tracking |
| Busy collection delays answering | Separate ingestion and query capacity | Predictable query service | More independently operated components |
| Top results repeat syndicated text | Canonicalization and story clustering | More diverse evidence | Imperfect duplicate/entity resolution |
| A summary claims population sentiment | Separate aggregate analytics from top-k synthesis | Honest denominators | Labeling, deduplication and coverage work |
| Expensive research exceeds three seconds | Explicit deeper-analysis class | Clear user expectation | Additional queue and pricing policy |
4. Detailed architecture
Read diagram source
flowchart TD
SRC[Licensed feeds and official filing sources] --> COL[Bounded collectors with cursors and gap detection]
COL --> RAW[(Immutable source versions)]
RAW --> OUT[Transactional ingestion jobs or replayable log]
OUT --> NORM[Parse, resolve entities and cluster duplicates]
NORM --> KW[(Keyword index)]
NORM --> EMB[Embedding workers]
EMB --> VEC[(Vector index)]
KW --> READY[Search-visibility and version readiness]
VEC --> READY
READY --> CAT[(Publication catalog and coverage status)]
USER[Question with time range] --> AUTH[Identity, source rights and query limits]
AUTH --> PLAN{Query class}
PLAN -->|Evidence search| HYB[Parallel exact and semantic retrieval]
KW --> HYB
VEC --> HYB
HYB --> FUSE[Deduplicate, fuse and rerank candidates]
CAT --> CHECK[Current-version and rights validation]
FUSE --> CHECK
CHECK --> GEN[Bounded answer with evidence references]
PLAN -->|Population statistic| AGG[Defined corpus and validated aggregate]
RAW --> AGG
AGG --> GEN
GEN --> VERIFY[Numeric, citation and support checks]
VERIFY --> ANSWER[Checked answer, coverage and timestamps]
CAT --> ANSWER
A practical stack could use Kafka for replay, PostgreSQL for publication metadata, object storage for source versions, and Elasticsearch plus Qdrant for retrieval. A single engine offering adequate lexical/vector search can reduce synchronization complexity. Choose two engines only when evaluated quality, scaling or operations justify the added coordination.
API and records
POST /answers
{query, entities?, event_time_from, event_time_to, mode: fast|research}
→ request_id, status, answer?, claims[], sources[], coverage, warnings[]
GET /sources/status
→ permitted source IDs, last successful collection, searchable coverage,
current incident and freshness measurements
| Record | Essential fields |
|---|---|
| Source version | Source/document IDs, revision, immutable object, published/event time, available/observed time, correction/deletion state |
| Ingestion job | Idempotency key, source cursor or sequence, retry state and original enqueue time |
| Passage | Source revision, span, text hash, entity IDs and source-rights scope |
| Index readiness | Revision, engine, visible/searchable state and verified time |
| Publication catalog | Current permitted revision and tombstone, with independent engine coverage |
| Answer trace | Query cutoff, evidence revisions, claims, model/prompt version, coverage, validation and usage |
Do not publish licensed full text to users entitled only to snippets. Authorization belongs in retrieval and artifact delivery, including citations and cached answers.
5. Ingestion correctness: replay does not make two databases atomic
Key updates by source plus document identity. Use a trusted revision/sequence where available; otherwise maintain a collector's ordered revision history and reconcile with the authoritative source. A wall-clock timestamp alone is not a reliable total order for concurrent or corrected updates.
- Persist the source version and an ingestion job durably, using an outbox transaction when the job originates from a metadata database.
- Build immutable passage IDs that include the document revision.
- Write lexical and vector entries with repeat-safe IDs; duplicate delivery must not create duplicate passages.
- Confirm actual search visibility, not merely write acknowledgement.
- Publish the revision for the query policy once its required engine paths are ready.
- Reject obsolete candidate revisions against the current catalog; retire old artifacts asynchronously.
- Advance a coverage checkpoint only past completed work; record explicit holes and retry them.
Kafka's transaction semantics coordinate Kafka records and offsets under the appropriate configuration. External index writes need their own cooperation, deduplication and recovery. Two consumers writing independently to Elasticsearch and Qdrant do not form an atomic transaction.
Elasticsearch refresh controls search visibility. refresh=wait_for can wait for a refresh; forcing an immediate refresh for every document can create costly small segments. Batch appropriately and measure end-to-end visibility. Similar acknowledgement-versus-readiness distinctions apply to other index engines.
A source correction invalidates the old claim before the replacement is ready when correctness requires it. Temporarily omit that document or use a verified fresh lexical-only revision; do not silently serve the superseded fact to keep recall high. Deletion and revoked source rights use an authoritative deny/tombstone check while physical cleanup proceeds.
6. Freshness, event time and honest coverage
| Time or checkpoint | What it answers | Common mistake |
|---|---|---|
| Source publication/event time | When did the event or publication occur? | Substituting crawler time for “last hour” |
| Source availability time | When could this service first retrieve the revision? | Assuming every source exposes this precisely |
| First observation time | When did our collector first see it? | Hiding a slow polling interval from freshness |
| Searchable time | When could queries retrieve the revision? | Recording only index write acknowledgement |
| Processing checkpoint | Which source sequence/cursor prefix is fully processed? | Using the newest document timestamp despite older gaps |
| Requested knowledge cutoff | What information may this historical answer use? | Including later corrections without labeling hindsight |
Where availability time is unavailable, state that the metric measures observation-to-searchability and report collection cadence/provider delay separately. The service cannot certify ingestion of updates it has never observed without provider sequence/cursor or equivalent completeness evidence.
A watermark in event-time processing is a progress estimate under assumptions about late events. A durable contiguous processing checkpoint is a different thing. Do not turn either into a claim that the entire web is current. Idle sources need collection heartbeats: “no recent articles” and “collector is broken” are different states.
Executable example: eligibility before ranking
The following illustrates an event-time window, current revision and rights check for a candidate. The adapters must obtain the catalog and permitted sources from trusted service state.
def eligible_candidate(doc, *, start, end, known_by,
current_revisions, permitted_sources):
if start >= end:
raise ValueError("Time window must be nonempty")
identity = (doc["source_id"], doc["document_id"])
return (
doc["source_id"] in permitted_sources
and not doc["deleted"]
and current_revisions.get(identity) == doc["revision"]
and start <= doc["event_time"] < end
and doc["observed_at"] <= known_by
)
Use timezone-aware timestamps with a defined interpretation of publication versus event time. For a historical answer, current_revisions must be the catalog as known at that cutoff; today's catalog would suppress an old revision that was valid then. Respect present access restrictions even for historical queries. An indexing timestamp filter or TTL cannot establish these conditions.
TTL limits retention and may reduce storage; it cannot make a delayed update arrive sooner. Store long-lived filings according to the product's retrieval/licensing needs instead of deleting everything after 24 hours.
7. Retrieval, sentiment and financial correctness
For exact identifiers, first resolve the intended company and ticker exchange context. Use metadata filters, lexical retrieval and semantic retrieval under the same rights/time constraints. Run appropriate paths in parallel, deduplicate by source revision and combine ranks with RRF. Rerank a bounded candidate set, then validate the selected evidence again before generation.
A reranker estimates relevance. It does not prove the source is true, that a claim is supported or that source permissions remain current.
“What is the sentiment around this company in the last hour?”
- Clarify the corpus: licensed news, a specified social feed, or both separately.
- Distinguish entity sentiment from overall article tone; an article can praise a supplier while criticizing the target company.
- Cluster copied/syndicated stories and define whether the unit is a post, unique story or author.
- For a distribution, classify the eligible population or a documented representative sample with a validated model.
- Report counts, excluded/uncertain labels, collection gaps and the measurement window.
- Use retrieved representative passages to explain the distribution, without presenting the top ten hits as the whole population.
For example, 60 negative, 30 neutral, 10 positive and 20 uncertain unique stories give 60% negative among 100 classified stories, but 50% among all 120 eligible stories. Neither is a share of people or proof of future price movement. Collection bias remains even with perfect classification.
For a financial number, validate entity, filing revision, concept, period, currency and unit. A matching numeral in an unrelated source passage is insufficient. Use deterministic calculations and show input references. Conflicts or insufficient evidence should produce a qualified result or abstention. See financial research design.
8. Three-second latency and failure behavior
An illustrative complete-answer budget is 100ms authentication/routing, 400ms retrieval, 300ms reranking, 1,700ms bounded generation, 300ms validation and 200ms headroom. The sum is three seconds; adding component p95 values does not prove the end-to-end p95.
GPT-6 Luna is a current small-model candidate for short grounded synthesis; compare alternatives on this corpus and full validation path. The model documentation is a capability reference, not a guarantee of the latency budget. Long reasoning, external browsing and multi-step analysis belong to the separately budgeted research path.
| Failure | Response policy | Tradeoff |
|---|---|---|
| Vector index lags | Fresh lexical-only retrieval with explicit coverage | Lower semantic recall |
| Both indexes exceed freshness target | Stale/search-only response clearly labeled or no current answer | Lower answer availability |
| Synthesis times out | Return retrieved evidence and status if product contract allows | No completed synthesis |
| Source is unavailable | State the missing source and preserved last coverage | Partial perspective |
| Numeric/support checks fail | Omit unsupported claims or abstain | Shorter or absent answer |
| Ingestion backlog grows | Prioritize critical licensed sources, add bounded capacity and replay | Other sources may miss objectives |
Streaming an unvalidated financial claim and retracting it later is a different product contract from returning only checked claims. Stream progress/source metadata early if useful; buffer claims until their required checks complete.
Cache keys include source rights, query/entity/time semantics, model/prompt versions and evidence revisions. A cached “last hour” answer needs a fixed evaluated interval and explicit as-of time. Short TTL helps limit staleness but does not replace invalidation on correction, deletion or rights changes.
9. Cost per answer, not a guessed model bill
At continuous 50,000 queries/hour over 30 days, traffic is 36M queries/month. A standard short-context GPT-6 Luna call with 2,000 uncached input tokens and 300 billed output tokens costs, at the reviewed $0.10/$0.50 per million token rates:
Model cost/query = (2,000 × $0.10 + 300 × $0.50) / 1,000,000
= $0.00035
Monthly generation = 36,000,000 × $0.00035 = $12,600
| Component | Illustrative monthly budget |
|---|---|
| Replayable ingestion | $2,500 |
| Vector search | $1,800 |
| Lexical search | $2,000 |
| Bounded generation | $12,600 |
| Reranking compute | $800 |
| Partial total | $19,700 |
Only the model rate is drawn from published API pricing; infrastructure numbers are allowances, not a sized vendor quote. The ingestion scenario's terabytes of vectors must be benchmarked against a real deployment before accepting that vector-search allowance. Include embeddings, feeds, raw storage, query servers, retries, observability and support. Billed reasoning tokens belong in output; cache writes, regional processing and other tiers can change the calculation.
At 90% successfully sourced answers, dividing this partial budget by 32.4M successes gives about $0.000608 per success. It is still a partial cost. Rejected and failed requests consume resources too.
10. Prove the design and close the interview
- Replay duplicate and out-of-order updates; the newest eligible revision must remain authoritative.
- Correct a filing after one index is updated; no answer may mix the obsolete claim into the new revision.
- Inject a dead-letter hole while newer events arrive; coverage must not advance past the unresolved gap without disclosing it.
- Revoke a feed entitlement; retrieval, cached answers and source delivery must enforce it.
- Test late events, daylight-saving boundaries, historical cutoffs and idle-feed heartbeats.
- Compare retrieval quality, source diversity and claim support on labeled queries; evaluate sentiment with a defined population and denominator.
- Burst ingestion and queries together, then fail each dependency while measuring freshness and complete-answer latency separately.
SEC EDGAR APIs offer timely submission/XBRL data as well as nightly bulk archives. Use the appropriate incremental source for the freshness target and follow fair-access guidance. A nightly bulk download alone cannot deliver five-minute updates.
Interview follow-ups
1. Does a time filter solve freshness? No. It limits which indexed events qualify. Collection delay, processing backlog, search visibility and corrections determine whether the necessary events are indexed at all.
2. Why not always add Kafka and two search engines? A small source set can use a simpler collector and one adequate engine. Replay and independent scaling can justify a log; improved retrieval can justify another engine. Each adds operational and consistency work.
3. What does Kafka exactly-once mean here? Its transaction support does not atomically publish a document into two external indexes. Use repeat-safe writes, versioned publication and reconciliation, with destination-specific guarantees.
4. Can the top ten results establish market sentiment? They can illustrate retrieved evidence, but ranking selects a biased subset. Define and evaluate the intended corpus or sample before reporting a population statistic.
5. What would you do when the three-second deadline is impossible? Return the agreed evidence-only/degraded outcome or offer the separately budgeted investigation mode. Do not hide a slow task behind a fast time-to-first-token measurement.
6. How do you know an “as of” timestamp is honest? It describes a measured source/index coverage boundary and any gaps. The maximum timestamp of one retrieved document cannot establish complete coverage.
60-second interview answer
I would start with durable source versions and measurable keyword retrieval, then add semantic search and short cited synthesis where they improve quality. The design separates event time, collection delay and search visibility, and handles corrections through versioned publication and authoritative candidate checks. I would expose source gaps, keep population sentiment separate from top-ranked examples, and use a bounded fast path with explicit degraded outcomes. Final validation covers evidence support, numeric meaning, freshness, latency and total cost per successful answer.
Remember: Collect → Publish versions → Retrieve → Verify → Disclose coverage.