System designby Learnastra

System-design interview · Extended interviews

Design a recommendation platform

By Anup Rai

Return a useful, permitted set of content by narrowing candidates, ranking a bounded set and learning from correctly interpreted feedback.

You will learn to

  • Build a working recommendation baseline before adding learned ranking.
  • Explain candidate retrieval, ranking, eligibility and feedback as distinct stages.
  • Size online work and defend model quality, cold starts and safe fallback.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Keyword search and vector retrieval · Caching: cache hits, misses, write policies and invalidation · Message queues, event logs, delivery guarantees, and backpressure · Design a feature-flag and configuration platform · Authentication, authorization, and tenant isolation

Workload and timing examples are interview assumptions.

01Define usefulness before choosing a model

A recommendation service selects a small set of items from a much larger catalog. For this interview, it returns twenty videos with display metadata. Video storage and playback are separate services. The objective is useful viewing and satisfaction, subject to safety, diversity and latency requirements; maximizing clicks alone can reward misleading thumbnails or repetitive content.

Ask what user outcome defines a useful recommendation, which eligibility rules are mandatory, and whether anonymous or opted-out viewers must still receive results. The worked design serves contextual recommendations to those viewers and treats safety and access as hard constraints.

Use viewer U7 opening the home page. The catalog contains ten million videos, but the response needs only twenty. U7 has recently watched cooking lessons and already completed item I12. The service must find plausible candidates, remove unavailable or inappropriate items, rank the remainder and return a varied set. Feedback later describes what U7 actually saw and watched.

The complete flow is request context → candidate retrieval → eligibility checks → ranking → diversity rules → response → feedback. Retrieval narrows the search space; ranking compares the selected candidates in more detail. These are different jobs. Begin with regional popularity so the service works before personalization, embeddings or a trained model exist. Learned components should improve a measured baseline rather than become unexplained prerequisites.

02Functional requirements

Agree on what the service must do before choosing its components.

  1. Serve a recommendation page. Return up to twenty permitted videos with display metadata for authenticated or anonymous viewers. If fewer eligible items exist, return fewer rather than violate a content rule.
  2. Respect viewer context. Apply language and region constraints, completed-item filtering and explicit dismissals; support cursor-based continuation bound to the viewer/session and feed context.
  3. Support personalization choices. Use permitted history for personalized retrieval/ranking and offer contextual popularity after opt-out or for a cold-start viewer.
  4. Capture useful feedback. Record actual visibility, clicks and watch activity with response and event identities. These observations support quality measurement and later ranking changes; video storage and playback are separate services.

03Non-functional requirements

Use these hypothetical requirements for the worked interview. Confirm the assumptions with the interviewer; the numerical targets require testing and are not measured results. For latency, p95 and p99 mean that 95% and 99% of measured delays, respectively, are no greater than the reported value.

  1. Workload. Plan for a ten-million-item catalog, 5,000 peak recommendation requests/s and 1,000/s on average. Bound the candidates scored for each page rather than examining the full catalog.
  2. Latency and availability. Target response p95 below 200 ms and 99.9% availability with the chosen candidate count and feature access pattern. Degraded personalization may fall back to eligible contextual results within that budget.
  3. Content eligibility. Check the current policy available at the defined serving boundary. A high score never overrides a hard rule; if eligibility cannot be established, omit the item or fail appropriately. Later removal cannot recall an already delivered response, so playback needs its own current checks.
  4. Recommendation quality. Evaluate useful watching and long-term satisfaction alongside clicks, abandonment, repetition and safety violations. Define experiment metrics and guardrails before a rollout: higher clicks with worse completion or more complaints may be a regression.
  5. Consent and privacy. Opt-out disables history-based retrieval/ranking and invalidates personalized continuation state. Limit personal-history retention and specify how deletion reaches derived features, including its propagation delay; do not let an old cached page silently continue withdrawn personalization.
  6. Trustworthy observations. Separate returned items from actual impressions, validate feedback and count a retried event once within the supported replay horizon. Feedback delay may make features older; it must not be hidden as current complete data.

04Serve a useful popularity list end to end

Maintain a catalog containing item ID, region availability, language, publication status, creator and basic popularity. A scheduled job computes popular items per region and language using recent verified activity. The serving API retrieves the top two hundred for U7’s context, checks current eligibility and removes completed or explicitly dismissed items.

Sort the remaining candidates by the baseline score, then select twenty while limiting repeated creators. Return their IDs, display metadata and a response ID. Persist or otherwise durably capture enough response context to interpret subsequent feedback: which items were returned, their positions and the ranking version. The UI reports actual visibility separately from the server’s decision to return an item.

The feedback endpoint accepts a stable event ID and verifies the viewer/session and response context. A durable queue retains events for aggregation. Retried copies of an event must not become several views or watch sessions.

This baseline needs an API, catalog, popularity job and feedback pipeline. It already handles cold-start users and can remain the fallback later. Its weakness is relevance: two people with different interests in the same region see similar candidates. That observed limitation, rather than fashion, motivates personalization.

Design diagramA complete popularity recommendation flow

The first useful product includes response identity and feedback.

A complete popularity recommendation flowThe first useful product includes response identity and feedback. viewer to api: Region, language and session; api to catalog: Read candidates and eligibility; api to viewer: Twenty items and response ID; viewer to events: Actual visible/watch events; events to job: Deduplicate and aggregate; job to catalog: Refresh popularityRegion, language and sessionRead candidates and eligibilityTwenty items and response IDActual visible/watch eventsDeduplicate and aggregateRefresh popularityCLIENTViewer U7SERVICERecommendation APISTORECatalog and popularlistsQUEUEDurable feedbackqueueWORKERPopularityaggregationsyncreturnasync
Read each connection in order
  1. syncRegion, language and sessionViewer U7 → Recommendation API
  2. syncRead candidates and eligibilityRecommendation API → Catalog and popular lists
  3. returnTwenty items and response IDRecommendation API → Viewer U7
  4. asyncActual visible/watch eventsViewer U7 → Durable feedback queue
  5. asyncDeduplicate and aggregateDurable feedback queue → Popularity aggregation
  6. asyncRefresh popularityPopularity aggregation → Catalog and popular lists

05Bound online work before adding model complexity

At five thousand peak requests/s, scoring every one of ten million items would require fifty billion item scores/s. Even a 50-microsecond score would consume roughly 2.5 million CPU-seconds each second. This arithmetic explains why candidate retrieval is necessary, not merely an optional optimization.

If retrieval narrows to two hundred candidates per request, the peak is one million scores/s. At the same illustrative cost, that is fifty fully busy CPU cores; at fifty-percent utilization, roughly one hundred before redundancy. Feature lookup, filtering and serialization still consume budget. Benchmark the whole path rather than treating this multiplication as a deployment guarantee.

Ten million 128-dimensional float 32 vectors contain about 5.12 GB of raw values. A searchable index needs additional memory for its data structures, metadata and replicas. An embedding is a learned numeric representation that places related items near one another; raw vector storage is not the total cost of searching them.

At one thousand average requests/s and twenty returned items, the service returns 1.728 billion item positions/day. If every position produced a 300-byte event, that is about 518 GB/day before replication. Actual visible impressions differ, so use this as a workload estimate, not a claim that every returned item was seen.

06Preserve response and event identity

Recommendation request

GET /recommendations
Request information Purpose
Authenticated context or anonymous session Identify the intended viewer context.
Region and language Apply the requested content constraints.
Page size Bound the number of returned recommendations.
Optional cursor Continue the intended session’s filtered feed state.

Recommendation response

Returned information Purpose
Response ID Identify the served result for later feedback.
Item IDs and positions Identify what was returned and where it appeared in the result.
Ranking version Identify the policy that produced the result.
Continuation token Bind the intended session, filters and feed state.

A continuation token must not let one user adopt another user’s personalized page.

Feedback event

POST /feedback
Event field Meaning
Event ID Stable identity of this observation.
Response ID and item The served result and item the observation concerns.
Event type Visibility, click and completed watch are different observations.
Client event time When the client observed the event; the service records receive time separately.
Relevant measurements For example, watch duration when applicable.

Validate feasible durations and ordering without assuming every disconnected client uploads immediately.

Catalog and feature records

Record Meaning
Catalog Items and their eligibility.
User features Permitted history summaries, such as recent topic interests.
Item features Quality, freshness and topic.
Feature definition Units, defaults and update age.

These definitions keep the ranker from mistaking milliseconds for seconds or missing data for a measured zero.

Keep model version and compatible feature definitions in a deployment bundle. The bundle is a named combination checked before rollout; naming it alone does not prove compatibility. Feedback names the served version so offline analysis can explain which policy produced an outcome.

07Add personalized sources without losing the fallback

A regional popularity list can miss niche cooking videos that match U7’s interests. Add several bounded candidate sources: recent videos from followed creators, matching topics, similar items and regional popularity. Each source returns a limited set. Merge duplicate item IDs, check basic eligibility and cap the combined pool before expensive feature enrichment.

Similarity retrieval can use an approximate nearest-neighbor index over item embeddings. Approximate means it trades exhaustive comparison for faster search and may miss some mathematically nearest items. The goal is a useful candidate pool, not a proof that a vector neighbor is the best video. Evaluate retrieval separately from the final ranker.

For example, gather at most one thousand candidates, then retain two hundred for richer scoring using a cheap preliminary score and source quotas. The exact limits are tunable budgets. A source that returns nothing should not stall the entire page; popularity remains available.

New users rely on declared interests and context. New items need metadata-based retrieval or a controlled exploration allocation because they lack watch history. Popularity-only feedback can otherwise keep them invisible forever. Exploration accepts some uncertainty to collect evidence, subject to the same safety and eligibility rules as ordinary recommendations.

08Rank a bounded set and apply product constraints

Fetch the two hundred candidates’ features in batches, not one network round trip per item. A simple explainable score can start with:

score = 0.60 × interest + 0.25 × quality + 0.15 × freshness

Each input is normalized to the intended zero-to-one scale. These weights are illustrative product choices, not universal constants.

Candidate example

Item Score Eligibility and selection
I11 0.60 × 0.9 + 0.25 × 0.8 + 0.15 × 0.6 = 0.83 A candidate for selection.
I12 0.825 U7 already completed it, so filtering removes it regardless of score.
I13 0.48 May still be selected if it is the next useful, allowed option.

A hard content rule is not just a small negative weight.

A trained ranker can replace the hand-set formula after offline evaluation and controlled online testing. It predicts an explicitly defined target from the same documented features. Missing features use trained or specified defaults; incompatible features trigger a fallback rather than arbitrary interpretation.

Finally apply creator caps and diversity rules to the ranked list. This may lower the raw predicted score while improving the overall experience. Select the page under those rules, even when it differs from the twenty highest scores.

09Keep training away from the request deadline

Scale the serving API horizontally and cache reusable item metadata and popularity lists. Personalized feature access uses bounded timeouts and privacy-aware keys. The online path retrieves candidates, batch-loads features, scores a bounded set, checks eligibility and returns the response. It does not train a model while U7 waits.

A separate pipeline consumes feedback, deduplicates stable event identities and updates aggregates. Scheduled training joins examples with feature values appropriate to the event’s historical context. Using information that became available only afterward leaks the future into evaluation and can make an ineffective model appear excellent offline.

Deploy a checked model/feature/index combination gradually. Canary traffic and an easy switch to the popularity baseline limit the impact of a bad release. An item index and model may have different update cadences, but their representation and schema compatibility must be explicit.

Keep queue lag and feature age visible. Durable events may arrive late, so freshness and completeness are different properties. If one retrieval source or feature service times out, use a documented simpler path within the response budget. The fallback still applies current eligibility checks; degraded relevance does not authorize unsafe content.

Design diagramBound the online path; learn asynchronously

Training and deployment improve the serving bundle without occupying the request deadline.

Bound the online path; learn asynchronouslyTraining and deployment improve the serving bundle without occupying the request deadline. api to sources: Retrieve plausible IDs; api to features: Batch facts and current policy; api to rank: Score bounded candidates; api to feedback: Response context and outcomes; feedback to train: Historical examples; train to rank: Checked gradual deploymentRetrieve plausible IDsBatch facts and current policyScore bounded candidatesResponse context andoutcomesHistorical examplesChecked gradual deploymentSERVICEOnline servingSERVICEBounded candidatesourcesSTOREFeature andeligibility storesSERVICECompatible rankerbundleQUEUEDurable feedbackWORKEROffline training andevaluationsyncasync
Read each connection in order
  1. syncRetrieve plausible IDsOnline serving → Bounded candidate sources
  2. syncBatch facts and current policyOnline serving → Feature and eligibility stores
  3. syncScore bounded candidatesOnline serving → Compatible ranker bundle
  4. asyncResponse context and outcomesOnline serving → Durable feedback
  5. asyncHistorical examplesDurable feedback → Offline training and evaluation
  6. asyncChecked gradual deploymentOffline training and evaluation → Compatible ranker bundle

10Interpret observations before learning from them

Returning I11 at position seven does not prove U7 saw it. The application may display only the first screen, the request may be abandoned or the item may be hidden. Record assignment, actual visibility, click and watch as distinct event types. This gives training and experiment analysis a chance to distinguish opportunity from outcome.

A retried visibility event reuses its original event ID. The consumer’s deduplication and aggregate update must be coupled, for example in one transactional sink, so a crash does not count the same observation twice. Retain deduplication state for the supported replay horizon; a key forgotten too early cannot protect a later retry.

An absent click is not automatically a negative preference when visibility is unknown. Likewise, a ten-second watch means different things for a twelve-second clip and a two-hour lecture. Define target labels with product context rather than training directly on whatever telemetry happens to be easiest to collect.

Use stable experiment assignment and record the policy actually served. Compare satisfaction, safety and latency guardrails in addition to the target metric. Position bias and exploration make causal conclusions harder; detailed counterfactual methods belong to advanced analysis, not an unsupported claim that raw clicks reveal pure preference.

11Recover serving and test the learning loop

Serving failure choices

Failure Behavior
Personalized retrieval fails Return eligible contextual popularity.
Rich features are missing Use a tested simpler score.
Current eligibility cannot be established Omit the candidate or use a verified eligible pool.
Feedback pipeline is delayed Serving may use older allowed features while exposing their age to operations.

Trace U7’s page through request, response ID, visible impression and watch event. Then repeat the feedback event, delay it, remove I11 before a later request and disable personalization. Verify one counted event, documented late-data behavior, no newly served removed item under the policy boundary and a contextual response after opt-out.

Measure candidate recall on judged examples, ranking quality, response latency, feature age, queue lag, duplicate rate, creator concentration, cold-start coverage and outcome guardrails. A good offline ranking metric does not by itself prove user benefit; validate with a controlled rollout.

Protect feedback APIs from fabricated events and limit retained personal history. Group monitoring metrics by bounded categories rather than creating a new metric series for every user. The next investment should follow evidence: poor retrieval needs better candidates, slow feature reads need serving work, and misleading labels need measurement repair rather than a larger model.

12Check the design against its requirements

Before closing, check the final design against the agreed requirements. FR means functional requirement and NFR means non-functional requirement; the numbers refer to the lists above. These are proposed validation checks, not test results.

Requirement Mechanism in the final design Validation and remaining limit
FR 1, 2; NFR 3 Bounded retrieval, current eligibility checks and diversity selection construct the page. Remove a high-scoring item, mark I12 completed and provide fewer than twenty eligible candidates. Return only allowed items; check playback separately.
FR 3; NFR 2, 5 Contextual popularity and consent-bound continuation provide a private fallback. Opt out while holding a personalized cursor and fail a candidate source. Require contextual allowed results, invalidated personalized state and documented deletion handling.
NFR 1, 2 Limited candidate sets, batch feature reads, replicas and bounded timeouts control online work. Load-test 5,000 requests/s, feature failures and hot contexts; measure p95 and availability. The 100-core arithmetic excludes additional serving work and is not a benchmark.
FR 4; NFR 6 Response context and transactional event deduplication preserve observation meaning. Return an item without displaying it, then replay or delay a real impression. Do not invent visibility or count the retry twice.
NFR 4 Judged retrieval evaluation and a controlled rollout compare against the popularity baseline. Measure usefulness, complaints, diversity and latency together. Improved offline scores or clicks alone do not prove user benefit.

13Rapid revision

Rehearse the numbered functional requirements and non-functional targets first. Use this table to recall the mechanisms, then close with the requirements check above.

Remember: Returned is not seen; observation drives learning.

Prompt Recall the mechanism and boundary
Why start with popularity? It provides measurable results without user history and remains a fallback
Why retrieve before ranking? Scoring all ten million items in detail for every request costs too much
What does retrieval produce? A limited set of plausible candidates that still need ranking
What do features need? Known units, missing-value defaults, update age and compatibility with the ranker
Why filter separately? A high score cannot make a forbidden item eligible
Why rerank after scoring? Limit repeated creators and vary the page’s content
Is a returned item an impression? Only actual visibility under the agreed measurement rule counts
What makes replay safe? Save the processed event identity together with its aggregate update
What prevents future leakage? Use only features available when the recommendation was made; later outcomes can supply labels
What survives a personalized outage? Return eligible contextual results within the response deadline

Close with: “I first serve a complete popularity feed and measure it. Candidate retrieval makes richer ranking affordable, while eligibility and diversity remain explicit product rules. A separate feedback and training pipeline improves relevance, and the original baseline remains an operational fallback.”

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

Why is regional popularity a valid starting design?

Reveal a model answer

It provides useful context-sensitive results without historical personalization or a learned model. The complete flow still filters eligibility, limits repeated creators, records response identity and collects actual feedback. It becomes both a comparison baseline and an outage fallback.

What the answer must demonstrate: It provides useful context-sensitive results without historical personalization or a learned model.

Applied · Question 2

Why not score all ten million videos on every request?

Reveal a model answer

At five thousand requests/s that is fifty billion scores/s. At the illustrative 50 microseconds each, it needs about 2.5 million busy cores before other work. Retrieval narrows the pool so richer ranking is affordable.

What the answer must demonstrate: At five thousand requests/s that is fifty billion scores/s.

Foundation · Question 3

What is the difference between candidate retrieval and ranking?

Reveal a model answer

Retrieval finds a limited set likely to contain useful items using sources such as followed creators, topics, similarity and popularity. Ranking spends more information and computation comparing that set. A perfect ranker cannot select a relevant item retrieval never supplied.

What the answer must demonstrate: Retrieval finds a limited set likely to contain useful items using sources such as followed creators, topics, similarity and popularity.

Applied · Question 4

I12 has a high score but U7 completed it. What happens?

Reveal a model answer

Under this product contract it is filtered out. Eligibility and completion rules are explicit constraints, not tiny penalties a sufficiently high engagement score can overcome. Diversity rules then shape the eligible ranked page.

What the answer must demonstrate: Under this product contract it is filtered out.

Applied · Question 5

Why is returning an item not enough to label it as ignored?

Reveal a model answer

The viewer may never have seen it. Record actual visibility separately from the returned list and use labels appropriate to the observation. Unknown exposure is not the same as a negative preference.

What the answer must demonstrate: The viewer may never have seen it.

Follow-up · Question 6

What is future leakage in this design?

Reveal a model answer

Training or evaluation uses facts that were not available at the historical decision, such as later popularity or a future watch. Offline metrics then overstate what serving could have achieved. Use historically appropriate feature values and availability.

What the answer must demonstrate: Training or evaluation uses facts that were not available at the historical decision, such as later popularity or a future watch.

Applied · Question 7

The personalized feature store times out. What should the API return?

Reveal a model answer

Use a tested simpler score or contextual popularity within the latency budget, while retaining eligibility checks. Expose feature age and fallback rate to operations. A relevance outage should not become a reason to serve forbidden content.

What the answer must demonstrate: Use a tested simpler score or contextual popularity within the latency budget, while retaining eligibility checks.

Follow-up · Question 8

What changes when U7 disables personalization?

Reveal a model answer

Stop using personal history for candidate retrieval and ranking, invalidate personalized continuation state and serve contextual results. Handle stored and derived history under the documented deletion policy and propagation limits.

What the answer must demonstrate: Stop using personal history for candidate retrieval and ranking, invalidate personalized continuation state and serve contextual results.

Blank-page exercise · 45 minutes

Build the answer yourself

Design a video recommendation homepage with twenty results, ten million catalog items, cold starts and a 200 ms latency goal.

  • Agree the numbered functional requirements and non-functional targets: page behavior, eligibility, consent, load, latency, availability and usefulness.
  • Draw the popularity baseline including feedback.
  • Calculate why exhaustive ranking is infeasible.
  • Add bounded retrieval and batch feature access.
  • Walk through ranking, filtering and diversity.
  • Validate recommendations, context, opt-out, feedback, quality, latency and fallback against the numbered requirements; name remaining evaluation and privacy-policy choices.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Design a recommendation platformI11 is returned at position seven, but U7 closes the page before seeing it. Is that an impression?Recall first, then reveal

No. Record actual visibility separately from returned items. Retrieve a bounded candidate set, filter and rank it, then learn asynchronously from what the viewer actually experienced.

Returned is not seen; observation drives learning.

Return to lesson
Design a recommendation platformDoes including an item in the response count as an impression?Recall first, then reveal

No. Count the agreed event showing that the item was actually visible to the user.

Returned is not seen

Return to lesson

Final revision

Summary and interview notes

Find and rank a limited set of permitted candidates. Record what users actually see and do, then use those observations in a separate process to improve future recommendations.

Remember these points

  • Agree the numbered functional requirements and non-functional targets before designing components; validate the final design against them.
  • Popularity supplies a complete baseline.
  • Retrieval controls online work.
  • Eligibility and diversity constrain ranking.
  • Actual visibility gives feedback its meaning.

Interview tips

  • Use a concrete candidate that is high scoring but ineligible.
  • Identify whether a problem lies in retrieval, ranking or measurement before proposing a larger model.

Important qualifications

  • Counterfactual evaluation and advanced exploration require additional statistical assumptions.

Continue after the core interview

Explore the advanced version

The advanced lesson keeps the full detailed design. Use these sections when you want to examine the stronger requirements and failure cases.

Technical references

Practice marks stay in this browser.