Interview problem: recommend eligible movies quickly, explain the supporting reason, and learn from user outcomes without exposing another person's viewing history.
A candidate generator narrows the catalog to a manageable set. A ranker scores those items for the product objective. Re-ranking applies list-level constraints such as diversity and eligibility. An explanation communicates an evidenced reason; fluent wording is not proof that the recommendation is useful or that the stated reason caused the model's choice.
This is a hypothetical Learnastra interview scenario with 50M registered users, 5M daily active users and ten recommendation sets per active user/day. Targets and cost allowances are assumptions.
1. Requirements and scope
Clarify the recommendation surface, number of visible items, regional rights, profile/household boundaries, age restrictions, personalization controls and what outcome the product values. A home-page list, “similar movies” panel and next-play suggestion can need different objectives.
Functional requirements
- Return a ranked set of movies eligible for the current profile, region and time.
- Combine useful collaborative, content, popularity and session signals.
- Provide a truthful reason for each result using verified catalog or permitted preference evidence.
- Support new users, new items and changing interests.
- Let users supply, correct or reset relevant preferences under the product's privacy controls.
- Record recommendation exposures and subsequent feedback for evaluation/training.
- Support controlled model, ranking and explanation experiments with rollback.
Nonfunctional requirements
- Target p95 complete responses under 200ms, including explanations, for the online surface.
- Plan for 50M recommendation responses/day and traffic bursts separately from registered-user count.
- Propose 99.9% monthly serving availability with eligible, nonpersonalized fallback lists.
- Enforce current content rights, age/policy restrictions and profile access.
- Keep private user/session data out of shared explanation caches.
- Bound retrieval candidates, ranking work and background explanation spend.
- Measure satisfaction, relevance, coverage, diversity, latency and cost rather than optimizing clicks alone.
Tip: Define the unit. A response containing ten movies is one recommendation-set request and ten item placements; it need not be ten online LLM calls.
2. Estimate the online and background loads
5M daily active users × 10 sets = 50M responses/day, averaging about 579/s. A 10× planning peak is about 5,790/s. At 120ms mean residence under that peak, roughly 695 requests are in flight under stable assumptions.
| Quantity | Illustrative calculation | Implication |
|---|---|---|
| Daily item placements | 50M responses × 10 items | 500M placements, not necessarily model calls |
| Long-term user vectors | 50M × 128 float32 dimensions | 25.6GB raw before keys/replicas/metadata |
| Item vectors for a 100,000-title catalog | 100,000 × 128 × 4 bytes | 51.2MB raw before index overhead |
| All ordered item pairs | Roughly 100,000² | Roughly 10B pairs; exhaustive explanation generation is wasteful |
| New validated explanation texts | Assumed 100,000/day | A separate measured cache-fill workload |
These dimensions/catalog size are planning choices. Full feature stores and event logs can dwarf vector storage. At 500M placements/day, a 100-byte event allowance alone implies 50GB/day before indexes and replication; log the fields the experiment/training design needs without indiscriminate personal data collection.
3. Start with an eligible popularity/content baseline
A baseline recommends available titles in the selected language/genre, removes explicit dislikes and uses verified template reasons. It can serve new profiles without an expensive personalization model.
Add collaborative retrieval and a learned ranker when they improve the defined user outcome. Google's recommendation overview describes the common candidate/scoring/re-ranking stages. LLMs can be evaluated for bounded ranking tasks, but this high-volume, short-deadline design keeps language generation outside the online critical path.
| Baseline flaw | Change | Benefit | Cost or limit |
|---|---|---|---|
| Same popular items for everyone | Collaborative and content candidate sources | More personal relevance | Behavioral data and cold-start issues |
| Long-term profile ignores tonight's goal | Session features or session candidates | Faster response to current intent | One unusual session can distort preferences |
| Highly similar items dominate | Diversity/coverage-aware re-ranking | More useful list variety | Possible relevance tradeoff |
| New items have no interactions | Content features and controlled exploration | Earlier discovery | Uncertain quality and exposure cost |
| Explanation cache miss blocks response | Verified template immediately, background richer text | Predictable miss-path latency | Simpler wording on misses |
| Click optimization harms satisfaction | Better objective and experiment guardrails | More appropriate user value | Delayed/noisy labels |
4. Detailed architecture and contracts
Read diagram source
flowchart TD
LOG[(Exposures, watches and feedback)] --> DATA[Time-correct training examples and features]
DATA --> TRAIN[Train and evaluate retrieval/ranking versions]
TRAIN --> UV[(Versioned user/item vectors and feature store)]
CAT[(Catalog facts and current eligibility)] --> CONTENT[(Content candidate index)]
REQ[Authenticated profile request] --> SCOPE[Resolve profile, region, policy and experiment]
SCOPE --> FETCH[Load permitted long-term and session features]
FETCH --> MULTI[Parallel collaborative, content and popularity candidates]
UV --> MULTI
CONTENT --> MULTI
MULTI --> FILTER[Merge IDs, deduplicate and filter eligibility]
CAT --> FILTER
FILTER --> RANK[Compact learned scoring]
RANK --> LIST[List-level diversity and final current eligibility]
LIST --> REASON[Verified reason codes and current-user evidence]
REASON --> CACHE[Bounded cache read or factual template]
CACHE --> OUT[Movies, reasons and response ID]
REASON -. bounded best-effort fill .-> BG[Generate generic wording from approved facts]
BG --> VALID[Validate claims, language and policy]
VALID --> RC[(Shared generic reason cache)]
RC --> CACHE
OUT --> EXP[Visible-exposure and outcome events]
EXP --> LOG
Candidate engines can run in parallel with individual deadlines. Final eligibility is applied after ranking as well because catalog rights can change while retrieval is underway. If rights cannot be verified, omit the item rather than treating an old cache entry as authorization.
GET /profiles/{id}/recommendations?surface=home
→ response_id, items[{movie_id, reason, reason_type}], model_versions
POST /recommendation-events
{event_id, response_id, item_id, event_type, occurred_at}
→ accepted|duplicate|invalid
The server resolves profile access and experiment assignment. Events are validated against issued responses and available authoritative playback signals; a browser-reported watch is not automatically reliable training truth.
| Record | Essential fields |
|---|---|
| Profile state | User/profile scope, explicit preferences, personalization policy and version |
| Interaction/exposure | Response, item, position, visible exposure, event/outcome time and policy/model version |
| Catalog version | Item facts, region/age/availability rules and effective times |
| Model bundle | Retrieval/ranker versions, compatible embedding space, feature schema and evaluation |
| Recommendation trace | Candidate sources, scores, applied constraints and eligible reason evidence |
| Generic explanation | Catalog/policy/language versions, item pair, reason type and validated wording |
| Personal explanation | Profile/permission scope, specific evidence and retention/version |
5. Candidate generation, matching spaces and ranking
Collaborative filtering uses patterns of user–item interaction. Content-based retrieval uses item attributes and the user's expressed or inferred interests. Both can contribute candidates; neither guarantees that a retrieved movie is eligible or desirable.
Matrix factorization jointly learns compatible user and item vectors. A separately trained text-embedding model defines another coordinate system. Equal dimension does not make those spaces compatible: do not query arbitrary text vectors with a matrix-factorization user vector. Search appropriate indexes separately and merge item IDs, or train a compatible two-tower retrieval model.
- Retrieve a bounded candidate allocation from collaborative, content, session and popularity sources.
- Deduplicate IDs and remove known ineligible or explicitly disliked items.
- Fetch point-in-time-consistent features and score the bounded set with a compact ranker.
- Apply list-level diversity, repetition/fatigue limits and product constraints.
- Recheck current eligibility and select a supported explanation reason.
- Return the list within the deadline; record which candidate paths timed out.
Dot product/cosine/index parameters must match the trained retrieval objective and vector normalization. A model rollout must pair query and item encoders with the same compatible index version. Use a versioned bundle/atomic routing switch or equivalent guarded rollout; a new user encoder against an incompatible old index can silently degrade recommendations.
The ranker can combine predicted satisfaction, completion, explicit feedback and business constraints. Avoid treating longer watch time as universally better: the objective should reflect the surface and user value. A title's length, popularity and existing exposure affect the labels.
6. Cold start and changing interests
There is no universal switch after ten watched items. Signal quality, sparsity and the user's current intent determine how much personalization is useful.
Read diagram source
flowchart LR
STATE[Current permitted profile and item evidence] --> NEW{Little reliable user history?}
NEW -->|Yes| BASE[Explicit preferences, eligible popularity and content]
NEW -->|No| MIX[Blend long-term and session candidates]
ITEM[New catalog item] --> CONTENT[Use content features and bounded exploration]
BASE --> RANK[Evaluate combined candidates]
MIX --> RANK
CONTENT --> RANK
RANK --> FEED[Observe exposure and outcomes]
FEED --> UPDATE[Update versioned features and review drift]
Keep long-term and session signals distinct. Blend them only in compatible learned spaces or combine separate candidate lists/scores using an evaluated procedure. “Recent watches count three times” is a tunable hypothesis, not a standard constant.
A watched title is not necessarily liked: autoplay, accidental starts and partial viewing matter. Explicit dislikes and “reset personalization” controls should affect serving promptly. New items can use metadata/content features before interactions accumulate; controlled exploration helps gather evidence but must respect eligibility and product constraints.
7. Explain a verified reason without inventing a story
Separate three claims:
| Claim | Necessary evidence |
|---|---|
| “This title shares the selected science-fiction genre.” | Verified item metadata and current explicit preference |
| “Because you watched this other title…” | This profile's permitted watch event and an actual related-item reason |
| “Because you enjoyed its director's style…” | Evidence of that preference; merely watching a movie is insufficient |
A catalog fact can be true while the implied account of the ranker's decision is false. Store valid reason codes from the selection logic and describe them accurately. If the system cannot support a personal rationale, show a generic factual attribute instead of inventing one.
For reusable wording, supply only approved item facts, language and reason category to a background model. Validate named entities, relationships and policy before caching. A fact sheet constrains available evidence but does not guarantee that a model follows it. Templates may be the more reliable choice for many reasons.
Fast cache miss with a safe background fill
def get_generic_reason(cache, jobs, catalog, request):
fields = ("catalog_version", "language", "policy_version",
"source_movie", "target_movie", "reason_type")
scope = {field: request[field] for field in fields}
key = tuple(scope[field] for field in fields)
try:
explanation = cache.get(key) # Adapter has a strict short deadline.
except (TimeoutError, ConnectionError):
explanation = None
if explanation is not None:
return explanation
fallback = catalog.verified_template(**scope)
# Local/bounded admission only; a full queue returns False immediately.
jobs.try_enqueue_nonblocking(key=key, scope=scope)
return fallback
The job adapter deduplicates concurrent fills, caps queue length and returns promptly without invoking a model. Queue saturation does not block the template response. Cache reads and catalog access also need bounded latency. A background provider outage must not become an online outage.
Share only generic catalog-based wording. Personal explanations require profile/audience scope, permission version and retention controls. Never put an entire user's viewing history into a shared item-pair cache. Language, catalog and policy changes invalidate or version the relevant entries; a TTL alone does not handle a corrected fact or revoked permission. See cache design.
8. Training and evaluation: exposure is part of the data
Record what was actually visible, item position, recommendation/experiment version and subsequent actions. A response sent to a client is not proof every item was seen. Unclicked or unexposed items are not automatically disliked.
Use time-based evaluation and features as they existed at prediction time. Prevent future watches, future catalog metadata and later feedback from leaking into historical examples. Group related profile/household data appropriately for the evaluation question.
Negative sampling selects comparison items during learning. Uniform, popularity-weighted and hard-negative choices affect the learned task and computational cost. Sampled ranking metrics depend on that candidate protocol; do not compare them directly to metrics over the full eligible catalog.
| Evaluation layer | Useful measurements | What it does not establish |
|---|---|---|
| Candidate retrieval | Recall@k on a defined eligible corpus | Final user satisfaction |
| Ranking | nDCG/ranking quality under fixed labels and candidate rules | Unbiased online effect by itself |
| Explanation | Factual support, faithful reason and user understanding | Recommendation quality alone |
| Online experiment | Predefined satisfaction/engagement outcomes and guardrails | Every long-term effect from a short test |
| Operations | End-to-end tails, failures, eligibility violations and cost | User value from uptime alone |
Recommendation exposure changes future interaction data, which can amplify popularity and narrow diversity. Controlled exploration and appropriately logged selection probabilities can support particular counterfactual analyses, but replay is not automatically unbiased. Define the estimator's assumptions and test its support rather than treating propensity logging as a cure for all selection bias.
For an A/B test, choose a stable assignment unit such as the profile or household according to the product's interference risks. Predefine outcomes, duration/stopping rules and guardrails. Monitor new users/items, language/region, diversity, satisfaction and latency. Randomly switching the algorithm on every request can contaminate a user's experience and make delayed outcomes hard to interpret.
9. Meet 200ms even when nothing is cached
| Online stage | Illustrative allocation |
|---|---|
| Identity, policy and request overhead | 20ms |
| Profile/session feature lookup | 5ms |
| Parallel candidate retrieval | 20ms |
| Compact ranking and list constraints | 50ms |
| Cached/template reason | 10ms |
| Serialization, network and headroom | 95ms |
| Total | 200ms |
This is a planning budget. Component p95 values cannot be summed to prove the full p95. A 95% explanation cache hit rate also does not prove the objective: hits can be slow elsewhere, and a large miss fraction can exceed the deadline if misses wait for generation. Measure complete responses under cold caches and dependency failures.
| Failure | Bounded response |
|---|---|
| User feature store unavailable | Eligible nonpersonalized fallback with generic reasons |
| One candidate engine times out | Continue with available candidate sources and record degraded coverage |
| Ranker fails | Known baseline ordering under the same eligibility rules |
| Explanation cache/model unavailable | Verified template; defer richer wording |
| Current rights cannot be established | Omit affected titles; do not bypass policy |
| Catalog/model versions disagree | Use a compatible prior bundle or fallback until rollout is corrected |
10. Economics and release plan
An online LLM call for each of 50M daily responses would cost $50,000/day at a hypothetical $0.001/call. That figure demonstrates the traffic multiplier; it is not a named provider quote.
Instead, assume 100,000 new generic explanation texts/day, each using 300 input and 60 billed output tokens. GPT-6 Luna standard short-context rates of $0.10/$0.50 per million yield $0.00006/text, or $6/day and $180 per 30-day month, before validation and infrastructure. API pricing.
| Serving component | Assumed cost per million responses | At 50M/day |
|---|---|---|
| Gateway, authorization and session lookup | $2 | $100 |
| Candidate retrieval/ANN | $8 | $400 |
| Ranker and list constraints | $12 | $600 |
| Feature/catalog/reason-cache reads | $3 | $150 |
| Telemetry and response/network overhead | $2 | $100 |
| Online serving subtotal | $27 | $1,350/day |
| Background explanation model | Separate 100,000-text workload | $6/day |
| Explanation validation/write allowance | Separate hypothetical allowance | $10/day |
| Partial machine/service total | $1,366/day |
The partial allocation is $40,980 per 30-day month or about $0.00002732/response. Add training, catalog ingestion, experiments, review and staffing where omitted. Do not count the same infrastructure twice. These tiny unit costs are credible only after real capacity benchmarks support the peak and tails; half the traffic may still require almost the same reserved infrastructure.
Roll out compatible feature/retrieval/ranker bundles through offline evaluation and controlled online traffic. Keep ranking quality and explanation truthfulness under separate ownership and metrics. Test cold caches, unavailable items, profile switches, deletion/reset requests, new-item bursts and incompatible model/index versions. Roll back the bundle without losing current eligibility or privacy restrictions.
Interview follow-ups
1. Why separate candidate generation from ranking? Deeply scoring every catalog item on every request is expensive. Candidate sources reduce the work; a stronger bounded ranker can then use richer features, followed by list and eligibility rules.
2. Can matrix-factorization user vectors search arbitrary text embeddings? No. Equal dimensions do not imply a shared coordinate system. Use compatible jointly learned representations or merge IDs from separate retrieval paths.
3. Does a new user need ten watches before personalization works? No universal threshold exists. Explicit preferences, content, popularity and sparse interaction signals can be blended and evaluated continuously.
4. What happens on an explanation cache miss? Return a verified template within the online deadline and admit a deduplicated background fill only if capacity is available. Do not wait for an LLM to preserve the appearance of personalization.
5. What makes an explanation unfaithful even if its facts are true? It may invent a preference or claim a reason that did not support selection. Keep reason evidence and distinguish a generic item attribute from a personal rationale.
6. Why can better offline nDCG accompany worse user satisfaction? The labels/objective may be misaligned, features may leak future information, the evaluation candidate set may be unrealistic, or exposure may create repetition and feedback effects. Inspect the protocol and online outcomes.
7. Is a shared movie-pair reason safe to reuse? Only if it contains generic permitted facts and matches the language/catalog/policy version. User-specific history or preference claims require that user's current evidence and audience scope.
60-second interview answer
I would separate candidate retrieval, ranking, list constraints and explanation. Compatible collaborative and content paths supply a bounded candidate set, while current rights and profile policy control what can be shown. Verified reason codes drive templates or validated background-generated wording, so a cache miss never waits for an LLM. Training records exposures and time-correct features, and experiments measure satisfaction and important slices. I would verify cold-start quality, truthful explanations, privacy, end-to-end latency and the complete serving cost.
Remember: Retrieve → Rank → Apply constraints → Explain from evidence → Measure outcomes.