System designby Learnastra

System-design interview · Extended interviews

Design a recommendation platform

By Anup Rai

Design candidate retrieval, feature serving, bounded ranking, eligibility, feedback quality, model bundles and controlled experiments.

You will learn to

  • Explain candidate generation, ranking, filtering, and diversification with actual item scores.
  • Construct training examples using the same feature definitions and values that were available when the live recommendation was made.
  • Design cold-start/failure fallbacks and experiments that measure more than clicks.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Keyword search and vector retrieval · Caching: cache hits, misses, write policies and invalidation · Message queues, event logs, delivery guarantees, and backpressure · Design a feature-flag and configuration platform · Authentication, authorization, and tenant isolation

Workload and timing examples are interview assumptions.

01Problem and scope

A recommendation service selects useful eligible items without an explicit search query. Define the objective first: this design targets useful viewing and satisfaction, while limiting repeated creators, unsuitable content and other agreed harms, not clicks or watch time alone. It returns up to twenty item IDs and display metadata from ten million items, with contextual/personalized modes, opt-out, pagination and consented feedback. Playback is separate. Add a model only if evaluation shows better viewing or satisfaction outcomes while meeting the latency and cost budgets.

Clarify the product objective

Candidate: “Is success a click, a completed watch, or a satisfied viewer?” Interviewer: “Useful viewing and satisfaction, with diversity and safety guardrails.” Candidate: “I will support anonymous/contextual and personalized home recommendations. A user can disable personalization. Deleted, blocked or region-restricted items must pass a final eligibility check, regardless of their score.” Clarify that a short suitable video can be useful without a long watch; maximizing total watch time alone is not our agreed goal.

Included and excluded surfaces

Include home-page recommendation, pagination, feedback, new users/items, and controlled model experiments. Exclude ad auctions, model architecture research, payments and playback delivery. We will begin with a popularity query. Add more advanced learning only when a controlled evaluation improves the agreed viewing and satisfaction measures without exceeding the serving limits. A neural model is not itself a product requirement.

02Functional requirements

  1. Get a page. Return distinct eligible items with stable identities and a stated fallback mode.
  2. Continue scrolling. Continue the same short-lived result session without repeating its previous items.
  3. Turn personalization off. New requests use contextual candidates and no personal behavioral features.
  4. Record feedback. Associate a visible impression, watch or dismissal with a valid item/page token.
  5. Publish an item. Admit it through catalog policy and give it a way to appear in recommendations before it has interaction history; this is the cold-start problem.
  6. Delete or block an item. Final eligibility checks exclude it according to the authoritative policy contract.

Result-page contract

User U27 requests up to twenty suggestions in the selected language and region. The service returns ordered item IDs, display metadata and an opaque page token. It can return fewer than twenty when eligible inventory is exhausted. An already-seen exclusion is a product policy: define whether it means a completed watch, any exposure, or an explicit dismissal. Here, exclude completed watches and dismissed items; returning an item that never appears on screen does not mean the user watched it.

Constraints and exclusions

An assignment to an experiment is not an exposure. A server response is not proof that an item became visible on a screen. Separate those events so a user who closes the app before rendering does not create twenty fabricated impressions. Similarly, missing feedback may mean lost telemetry rather than a negative preference. State these distinctions before choosing the event pipeline.

03Non-functional requirements

  1. Workload assumption. 1,000 average and 5,000 peak page requests/s, ten million active catalog items, and twenty desired results/page.
  2. Latency and availability. Server-side p95 below 200 ms, p99 below 400 ms and 99.95% successful eligible requests/month. Network and playback startup have separate budgets.
  3. Success definition. A degraded contextual response may count as success; a response containing forbidden items does not.
  4. Freshness. Popularity/features may lag a few minutes. Current eligibility cannot silently fall back to obsolete permission caches.
  5. Quality. Optimize useful viewing with satisfaction, repeated-exposure, creator-diversity and content-quality guardrails. Agree on metric definitions before claiming a model is better.
  6. Retention and durability. Retain identifiable feedback for an illustrative thirty days; apply an explicit deletion/consent policy to derived data. Accepted feedback survives one event-store node failure; unacknowledged client events remain retryable.
  7. Serving continuity. Training downtime can delay improvements without stopping online recommendations. Every request uses one compatible model/feature/index bundle.

Final authorization boundary

Situation Required behavior
Final policy read starts after an acknowledged deletion or consent change Observe that change
Consent changed Invalidate the older personalized session as well as checking item eligibility
Earlier-authorized response arrives later Disclose the in-flight limit; an already rendered screen cannot be erased
Playback starts Authorize access separately
Policy authority unavailable Serve independently permitted public fallbacks under their contract, or fail clearly

An empty result is preferable to invented authorization. These are illustrative service targets, not current measurements from a named product.

04Capacity estimates

Retrieval chooses a manageable candidate set using relatively cheap indexed lookups or similarity search. Ranking then spends more work comparing those candidates with richer inputs. Keeping these stages separate is what makes a large catalog affordable to serve: the detailed scorer does not examine every item for every page.

At peak, scoring all ten million items for every request would require fifty billion scores/s. With an assumed 50 microseconds of CPU per rich score, that is 2.5 million CPU-seconds every second before feature I/O. The number demonstrates why cheap retrieval precedes careful ranking. It is not a measured benchmark of a named model.

Resource Calculation Architectural consequence
Rich scoring after retrieval 5,000 requests/s × 200 items = 1M scores/s Batch features and score only a bounded shortlist
Ranking CPU 1M × 50 μs = 50 busy cores At 50% target utilization, budget about 100 cores before redundancy/skew
Raw item vectors 10M × 128 dimensions × 4 B = 5.12 GB Approximate nearest-neighbor (ANN) index structures, metadata and replicas require extra RAM
Peak visible item events 5,000 pages/s × 20 = 100,000/s upper bound Batch events; distinguish returned from actually visible
Daily events at average load 1,000 × 86,400 × 20 = 1.728B At 300 B/event, about 518 GB/day before compression/replicas
Thirty-day raw retention 518 GB × 30 ≈ 15.6 TB Event retention and training reads are material cost drivers

The item-event estimate is an upper bound if every result is visible once. Page envelopes and item ordinals reduce repeated metadata. For serving, allocate context 20 ms, retrieval 40, features 30, ranking 60, final eligibility 20 and response/overhead 30. These are per-stage time budgets. Adding separately measured stage p95 values does not establish the p95 of complete requests, because the slowest requests can differ between stages. Measure complete requests under correlated slowdowns. If retrieval overruns, cancel it and preserve time for final filtering instead of spending the safety budget on one more candidate source.

05APIs and contracts

Request and authenticated page token

POST /v1/recommendations receives {surface:"home",count:20,sessionId:"S8",cursor:null} under authenticated user U27. Region, age policy and consent come from trusted account/request context, not arbitrary client claims. Response R81 contains {items:[I11,I13,...],pageToken:X9,nextCursor:C2,bundle:B7,mode:"personalized"}. The page token is opaque or authenticated; it binds the result IDs/positions, session, expiry and permitted feedback identity without exposing personal features.

Continuation, consent and replay

  • Continue the cursor. A continuation uses C2 to retrieve the next slice of the short-lived ranked session.
  • Retain session order. Store its ordered candidate IDs and already-delivered offset, then recheck current consent and item policy before each page.
  • Recheck consent version. Store the session's consent version; a changed version invalidates its old personalized ordering and requires recomputation under current consent.
  • Allow a shorter page. Filling holes can produce a shorter page; it must not resurrect an ineligible item to preserve pagination shape.
  • Expire explicitly. An expired cursor returns a clear restart instruction.
  • Choose page-replay semantics. A request ID correlates attempts; a recommendation GET/POST need not reserve a financial-style business operation, but a session-page number can replay a stored page if stable retry ordering is desired.

Durable feedback acceptance

POST /v1/events accepts bounded batches such as {eventId:E55,pageToken:X9,item:I11,kind:"visible",ordinal:0,clientTime:T}. Return 202 only after durable event-log acknowledgement. A duplicate E55 is deduplicated for downstream effects. Reject an item not present in X9, an expired/forged token, excessive batch sizes and unauthorized identities. Watch duration is bounded and validated; the client is not a trusted source of monetary or security facts. Rate-limit bots separately from ordinary retries.

06Data model and access patterns

Catalog and consent authority

Catalog authority stores Item(itemId, creatorId, language, regionPolicy, publicationState, policyVersion), indexed by item ID and by eligible language/topic for the baseline. Consent authority stores (userId, personalizationAllowed, consentVersion). Check current eligibility through a batch policy lookup. Candidate and feature stores may lag behind that source of truth, so they cannot grant access.

A serving bundle names the artifacts that must work together: model M7, feature schema F4, retrieval index I7, and the corresponding transforms and defaults. A feature schema specifies what inputs mean, including their units and representation. The bundle gives deployment checks one compatible set to validate before activation, so the model is paired with the inputs and retrieval index it expects.

Derived and serving records

Record Key and query Meaning
Candidate list (region,language,topic,indexVersion) Bounded ordered IDs from a named retrieval source
User/item feature (schemaVersion,entityId,featureName) Value, unit, event time and availability time
Model bundle bundleId=B7 Immutable manifest for M7, F4, I7, defaults and checksums
Result session (U27,S8,pageNumber) Ordered remaining IDs, consent version and served-page identity, with short TTL
Response intent X9 Returned item positions, bundle/experiment and request context
Outcome event eventId=E55 Validated visible/watch/dismiss event referencing X9 and an item

Feature definition and worked score

A feature is a measurable input, such as minutes watched in a topic during a specified past window. “Affinity” is not a self-explanatory database column: define its range, aggregation, time window, default and consent requirements.

For a simple worked example, give each item three normalized inputs between 0 and 1: topic affinity a, content quality q, and freshness f. Higher values mean a stronger signal. The sample values below are assumed inputs, not measured production features. Compute score = 0.6a + 0.25q + 0.15f. A deployed feature definition must additionally specify its source, time window, normalization and missing-value default.

Item Affinity a Quality q Freshness f Score Eligibility/result
I11 .9 .8 .6 .83 Eligible
I12 .8 .9 .8 .825 Excluded: completed watch
I13 .2 .9 .9 .48 Can be shown despite its lower score

Exclude I12 because it was completed; I13 can still be shown despite its lower score. These weights are illustrative and interpretable, not asserted production parameters.

07Basic working design

Indexed contextual popularity

Start with one stateless API and a relational catalog/event database.

  1. Build contextual popularity. A periodic job calculates recent popularity by region/language from validated outcomes.
  2. Read and filter bounded candidates. For user U27's request, read the top 200 candidates using an index on (region, language, popularity DESC, itemId), remove completed/dismissed and forbidden items, enforce a simple creator cap, and return twenty.
  3. Commit response intent before feedback. Persist response intent X9 before returning it, then ingest visible events independently.
  4. Measure the minimal product. Record page latency, eligible results returned, actual visible items, and the agreed viewing and satisfaction outcomes. Use these measurements as the baseline for later model changes.

Durable feedback

The database transaction on feedback inserts E55 under a unique key and records its accepted state. A lost response causes a same-ID retry, not another popularity increment. Popularity is periodically rebuilt from events, so a crash between accepting E55 and updating the derived score is repairable. Catalog deletion changes authoritative policy immediately; the popularity job may leave the stale candidate ID around, but the final check excludes it.

Result sessions and current eligibility

The result session stores a bounded candidate order for a few minutes. That trades modest memory for stable pagination; restarting on expiry is acceptable for this home feed. The baseline is easy to inspect: explain precisely why I11 was returned and what E55 changed. It is less personalized than later versions, but adding an opaque model before collecting reliable observations would make failures harder to diagnose rather than improve the contract.

architecture · baselineIndexed popularity with a real visibility event

Catalog policy is authoritative. Returned page X9 and visible event E55 are distinct facts; the popularity aggregate is derived.

Indexed popularity with a real visibility eventCatalog policy is authoritative. Returned page X9 and visible event E55 are distinct facts; the popularity aggregate is derived. client to api: 1. Request page R81; api to db: 2. Candidates, policy; save X9; api to client: 3. Return IDs and page token; client to api: 4. E55 when I11 is visible; api to db: 5. Insert E55 once; db to job: 6. Read accepted outcomes; job to db: 7. Rebuild popularity1. Request page R812. Candidates, policy; save X93. Return IDs and page token4. E55 when I11 is visible5. Insert E55 once6. Read accepted outcomes7. Rebuild popularityACTORUser U27’sapplicationSERVICERecommendation APISTORECatalog and event DBWORKERPopularity buildersyncasync
Read each connection in order
  1. sync1. Request page R81User U27’s application → Recommendation API
  2. sync2. Candidates, policy; save X9Recommendation API → Catalog and event DB
  3. sync3. Return IDs and page tokenRecommendation API → User U27’s application
  4. sync4. E55 when I11 is visibleUser U27’s application → Recommendation API
  5. sync5. Insert E55 onceRecommendation API → Catalog and event DB
  6. async6. Read accepted outcomesCatalog and event DB → Popularity builder
  7. async7. Rebuild popularityPopularity builder → Catalog and event DB

08Find the baseline flaws

Failure test What breaks and what must follow
Query/scoring fanout Suppose the baseline handles 500 page queries/s at the target tail latency. Peak demand is 5,000/s, so connection queuing rapidly consumes the 200 ms budget. Fetching rich item features with 200 serial lookups at even 1 ms each would exhaust the entire budget before ranking. More API instances cannot remove a shared database query bottleneck or serialize two hundred network round trips any faster.
Popularity and cold start Popularity also has a product flaw. A new high-quality woodworking item has no watch history, so it never enters the top 200. Without any exposure it cannot acquire the history needed to enter them. This feedback loop needs a discovery policy, not a faster cache. A creator cap prevents an endless row from one creator but does not itself solve cold start or topic coverage.
False exposure and incompatible features Now break measurement: the API returns twenty items, user U27 closes the app, and the server logs twenty impressions. Training treats their absent clicks as negatives even though none was visible. Another bug joins yesterday's exposures to today's popularity, allowing the model to learn from information it could not have known. Both inflate or corrupt evaluation. Finally, change a feature from seconds to minutes without changing its name: an old model can silently receive values sixty times smaller. We need explicit event semantics, temporal joins and compatible serving versions alongside capacity changes.

09Improve the design, step by step

1. Precompute candidate pools and isolate serving reads

  • Trigger: Repeated region/language queries motivate cached bounded lists generated from durable events.
  • Mechanism: APIs batch candidate/item reads and use stateless replicas. This removes expensive repeated aggregation and reduces database contention.
  • Benefit, cost and alternative: Costs are cache RAM, refresh lag and invalidation/rebuild operations. The new risk is a stale deleted candidate, so final eligibility stays outside this cache. Keep indexed database queries while they meet measured load; a cache is not mandatory merely because the catalog is large.

2. Combine several retrieval sources

  • Trigger: Popularity keeps showing established items, leaving new items and niche interests with little exposure.
  • Mechanism and tradeoff: Combine followed creators, content similarity, topic lists and controlled exploration. A learned embedding is a vector used to retrieve nearby items; an approximate nearest-neighbor index trades retrieval exactness for bounded query work. Merge perhaps 1,000 candidates, deduplicate, then take 200 to richer ranking. The benefit is broader recall; costs include indexes, freshness pipelines and relevance tuning. The new risk is that one source dominates or a retrieval filter removes the best item before ranking. A simpler topic/popularity union is preferable until a learned retriever improves measured coverage. The historical two-stage research example is in the technical references; our numbers are our exercise assumptions.

3. Add batched features and a versioned ranker

A feature is an input to the scoring model, such as recent topic watch time. Materializing a feature means computing and storing that value ahead of a request, so serving can read it without repeating the aggregation.

  • Trigger: The 200 serial lookups motivate multi-get by shard and one bounded model call.
  • Mechanism: Reranking applies creator/topic diversity after scoring. This provides more useful personalization while keeping the expensive stage small.
  • Benefit, cost and alternative: Costs are feature materialization, model CPU, defaults and deployment complexity. The new risk is training-serving skew, so one immutable bundle pins model, schema, transforms and retrieval compatibility. Use the interpretable weighted score or the baseline if the learned model's incremental benefit fails to justify those costs.

Training-serving skew means the model sees different input meanings or calculations during training and live use. The seconds-to-minutes mistake is one example; different missing-value defaults or using later information during training are others. Versioning the input definitions and reproducing what was available at the original decision address different parts of that mismatch.

4. Separate feedback learning from online serving

  • Trigger: The event volume and response-versus-visibility bug motivate durable ingestion, deduplication, point-in-time training datasets and controlled experiments.
  • Mechanism: Training can retry large jobs without holding an interactive request.
  • Benefit, cost and alternative: Costs are retention, joins, experiment infrastructure and delayed learning. The new risks are missing/late events and biased exposure. Track coverage by client/version and avoid treating unobserved events as known negatives. Real-time model mutation on every click is rejected here because it adds unstable feedback and rollout complexity without a stated freshness need.
Candidate method Mechanism Main limitation
Contextual popularity Aggregate eligible outcomes by language/region/topic Feedback can concentrate exposure on established items
Content-based similarity Match item metadata or content embeddings to declared interests or eligible history Similar content can become repetitive
Collaborative filtering Learn patterns from users' item interactions, such as item co-consumption or latent factors Sparse/new users/items and exposure bias limit evidence
Two-tower retrieval Encode request/user context and items separately; index item vectors and search with the request vector Query/item encoders and index must belong to a compatible model generation

For the learned option, precompute item embeddings offline and query the ANN index with the compatible request tower online. A different query encoder with the same output dimension is not necessarily in the same vector space. Rank the resulting small set with richer interaction features. These approaches can contribute candidates together; none removes consent, cold-start exploration or final eligibility.

10Detailed architecture

Bounded serving path

The API obtains trusted context, a deadline and one active bundle. A candidate coordinator queries independently bounded retrieval sources in parallel. Each returns IDs, scores/provenance and index version. The feature service batch-loads compatible user/item inputs; the ranker scores the shortlist; a final assembler rechecks eligibility and diversifies before storing X9 and responding. These are logical responsibilities: early deployments can place several in one process while retaining the same contracts.

Policy versus derived features

Catalog/consent authority owns permission and publication facts. Feature stores, vector indexes, popular lists and cached result sessions are derived. A stale index can delay discovery of a new item, but cannot make a deleted item eligible. Check the selected items in one policy batch rather than twenty network round trips. Permission changes and playback checks use the same authoritative policy semantics.

Feedback and bundle publication

The event collector validates and durably appends outcomes. Stream processors create fresh features and aggregates; a retained event lake supplies training with time-correct examples. Offline trainers publish validated artifacts to immutable storage and a bundle registry. Serving replicas warm the complete bundle before atomically activating its pointer. They continue with their pinned bundle or a declared baseline if the control plane fails; they never assemble an accidental mixture of files from successive rollouts.

Independent capacity and request deadline

Capacity scales independently across request serving, retrieval, scoring and learning. Each has admission limits. The queue/lake are not in the ranker's synchronous request path, although storing response identity must complete under its small allocated budget if we promise its durability. If that intent store is unavailable, either fail the measured/experiment path or explicitly mark an untracked baseline response; do not silently include it as a complete experiment observation.

architecture · finalBounded online serving and a versioned learning path

The serving group returns a page within its deadline; the policy group checks eligibility; the learning group processes observed outcomes. Publishing a model bundle changes scoring, not a user’s access rights.

Bounded online serving and a versioned learning pathThe serving group returns a page within its deadline; the policy group checks eligibility; the learning group processes observed outcomes. Publishing a model bundle changes scoring, not a user’s access rights. client to api: 1. Request R81; api to retrieve: 2. Trusted context and B7; retrieve to indexes: 3. Parallel candidate lookup; retrieve to rank: 4. Shortlist 200 IDs; features to rank: 5. Batch F4 inputs; rank to assemble: 6. Scores and B7 provenance; assemble to authority: 7. Current consent / exact item policy; assemble to pages: 8. Save X9 and cursor; assemble to client: 9. Eligible ordered page; client to events: 10. Actual visible/watch events; events to log: 11. Durable accepted E55; log to train: 12. Time-correct examples; train to features: 13. Versioned feature updates; train to bundle: 14. Validate and publish B8; bundle to api: 15. Activate complete bundle1. Request R812. Trusted context and B73. Parallel candidate lookup4. Shortlist 200 IDs5. Batch F4 inputs6. Scores and B7 provenance7. Current consent / exact itempolicy8. Save X9 and cursor9. Eligible ordered page10. Actual visible/watch events11. Durable accepted E5512. Time-correct examples13. Versioned feature updates14. Validate and publish B815. Activate complete bundleACTORUser U27’sapplicationSERVICEContext and page APIG1SERVICECandidatecoordinatorG1STOREVersioned candidateindexesG1STORENamespaced featurestoreG1SERVICEPinned model rankerG1SERVICEPolicy and pageassemblerG1STORECatalog / consentauthorityG2STOREResponse andsession storeG1SERVICEValidated eventcollectorG3QUEUEReplicated event log/ lakeG3WORKERFeatures and trainingjobsG3STOREImmutable bundleregistryG3syncasynccontrolG1 Serving plane / compatible request bundleG2 Current eligibility authorityG3 Feedback and model control plane
Read each connection in order
  1. sync1. Request R81User U27’s application → Context and page API
  2. sync2. Trusted context and B7Context and page API → Candidate coordinator
  3. sync3. Parallel candidate lookupCandidate coordinator → Versioned candidate indexes
  4. sync4. Shortlist 200 IDsCandidate coordinator → Pinned model ranker
  5. sync5. Batch F4 inputsNamespaced feature store → Pinned model ranker
  6. sync6. Scores and B7 provenancePinned model ranker → Policy and page assembler
  7. sync7. Current consent / exact item policyPolicy and page assembler → Catalog / consent authority
  8. sync8. Save X9 and cursorPolicy and page assembler → Response and session store
  9. sync9. Eligible ordered pagePolicy and page assembler → User U27’s application
  10. async10. Actual visible/watch eventsUser U27’s application → Validated event collector
  11. async11. Durable accepted E55Validated event collector → Replicated event log / lake
  12. async12. Time-correct examplesReplicated event log / lake → Features and training jobs
  13. async13. Versioned feature updatesFeatures and training jobs → Namespaced feature store
  14. control14. Validate and publish B8Features and training jobs → Immutable bundle registry
  15. control15. Activate complete bundleImmutable bundle registry → Context and page API

11Write path and acknowledgement

Validate and deduplicate feedback before feature or training updates. Publish compatible immutable model, transform, feature and index bundles after evaluation.

Numbered feedback and learning flow

  1. Commit response intent. The server commits response intent X9 for R81: returned IDs/positions, B7, experiment assignment and compatible feature/provenance metadata. A response intent is not yet a visible impression.
  2. Validate and durably accept actual visibility. User U27's client renders I11 and emits visible event E55 with X9 and its ordinal. The collector validates token, identity, item membership, size and event type, then appends to replicated durable storage before 202. The client retains E55 for bounded same-ID retry after an unknown outcome.
  3. Deduplicate the sink effect. Consumers deduplicate E55. An aggregate sink uses a unique processed-event marker and its counter change in the same local transaction, or builds immutable partition outputs that are atomically replaced. “At-least-once queue” alone cannot prevent a double counter increment.
  4. Wait for mature outcome labels. A watch event E56 references the same exposure. The learning pipeline waits an agreed period for outcomes such as a completed watch before labeling the example; this is the label-maturity window. Late events revise that example under a versioned policy. No click after a fully observed window is different from a missing visibility event.
  5. Construct point-in-time training data. Training reconstructs features whose event time and availability time are both no later than R81's decision time. Future information stays out. Deletion/consent filters apply to dataset creation and eligible serving features; retained identifiers permit required removal rather than leaving untraceable personal copies.
  6. Validate and canary a complete bundle. Validate an immutable new bundle B8, publish checksums/schema contracts, warm replicas, then canary it under a recorded experiment. Offline metrics can reject a bad candidate but cannot prove causal online benefit.

Sink commit versus offset acknowledgement

The sink is the destination that stores an event’s effect, such as a popularity counter. The queue offset records how far the consumer has processed. Saving the counter and acknowledging the queue are separate operations, so a crash between them can deliver the event again.

A processor crash after sink commit but before offset acknowledgement causes replay. The processed-event key returns the existing effect, so the event's contribution remains one. Monitor duplicate and invalid-event rates; a globally unique-looking ID is not proof that a client supplied truthful behavior.

Scoped identity and consent-aware processing

Event identity is scoped to the authenticated application/tenant and actor as well as eventId; store a payload fingerprint so a retry with changed page/item/type fails. A valid old page token proves prior presentation context, not present consent to continue personal learning. Check current consent at collection and at feature/dataset use under the declared policy; after opt-out, do not turn delayed personalized events into fresh behavioral features. Keep any operational aggregate telemetry only under its separate non-personal retention contract.

12Read and delivery path

Each request pins one serving bundle, retrieves bounded candidates and ranks eligible content. Current catalog policy and consent govern released results.

Numbered recommendation request

  1. Pin context and a compatible bundle. Authenticate U27, derive current consent/region and set an absolute request deadline. Read active bundle B7 once and retain a reference for R81. For personalization disabled, omit behavioral user features entirely and use contextual retrieval.
  2. Retrieve bounded candidates in parallel. In parallel, retrieve candidates from topic/popularity, followed creators and I7 similarity. Impose per-source deadlines and quotas; merge at most 1,000 unique IDs with provenance. Cold-start users use contextual pools; new items receive controlled exploration if policy permits.
  3. Batch compatible features. Select 200 candidates for scoring and batch-load F4 feature values. Validate schema, defaults and freshness. A timed-out shard supplies explicitly allowed defaults or causes a bounded baseline path, never values from an incompatible F5 namespace.
  4. Score and apply diversity. Apply M7 to these inputs. In the worked example I11 scores .83, I12 .825 and I13 .48. Discard I12 under the completed-watch rule even though it nearly outranks I11. Creator/topic caps may move other eligible lower-scored items ahead to improve diversity.
  5. Recheck consent and item policy. Perform the final authoritative policy read for both current consent and item eligibility. Compare the current consent version with the context and saved session. If it changed, discard the stale personalized ordering and features; when consent is now off, rebuild a contextual page within the remaining deadline or return a clear retry. This also applies to cached continuations, which cannot keep a personalized order merely by removing forbidden items. Remove ineligible IDs and backfill from already-scored permitted candidates while time remains. Return fewer results rather than bypass either check.
  6. Commit page identity and return. Commit X9, return the page/cursor and record latency/version/fallback diagnostics. Visibility/outcome events arrive later. The app must not replay an expired page token indefinitely as if a stale authorization were current.

Predetermined fallback

Ranking timeout invokes a predetermined contextual ordering with current policy checks. Cancel outstanding expensive work when its result can no longer meet the request deadline. Otherwise “fallback” returns quickly while abandoned ranker calls continue consuming enough compute to cause the next outage.

Bind final authorization to display bytes

Check permission for the exact title, thumbnail and other display fields that will be returned. Load them in a batch with immutable item-version references, then ask the policy authority to check those versions, their policy revisions and current consent. If a loaded version and policy reply disagree, reload and recheck within the deadline or omit the item. Render the checked representation; authorizing an older public version does not permit fetching a newer private title or thumbnail. A private-thumbnail endpoint also checks its own grant and current-access policy. These checks govern release of the response; they cannot recall an answer already authorized and sent.

13Correctness deep dive

Pin one compatible bundle

Bundle B7 means model M7, feature schema F4 in minutes, index I7 and transform/default versions D4. B8 changes the model and feature representation. Both remain immutable after publication; the online store keeps distinct schema namespaces. Activate the bundle by switching one pointer, so a request cannot pick up independently changed artifacts.

Publication and request protocol

The activation step uses compare-and-swap: replace the active bundle only if it still equals the expected previous bundle. A request then keeps a reference to the bundle it acquired, so a concurrent activation changes later requests without replacing this request’s model or feature definitions halfway through.

prepare(B8):
  verify artifact checksums, schema IDs and index dimension
  load M8; warm I8; confirm F5 availability or valid defaults
  run known-input predictions and policy/fallback smoke checks
  mark this serving replica READY(B8)
activate_on_replica(expected=B7, next=B8):
  require READY(B8)
  compare-and-swap active_bundle B7 -> B8
serve(R81):
  b = acquire_reference(active_bundle)
  candidates = retrieve(index=b.index)
  f = batch_features(schema=b.schema, candidates)
  require f.schema == b.schema and b.model.accepts(f.schema)
  score(b.model, f); final_policy_check(); release_reference(b)

Concurrent bundle switch timeline

Point-in-time training proof

Compatibility does not prove quality

These protocols establish compatibility and faithful examples; they do not prove the model improves satisfaction. That remains an experiment question. A perfectly versioned harmful objective is still a poor recommender.

sequence · bundle-raceR81 keeps B7 while R82 starts on B8

The request retains one immutable bundle. Feature namespaces and reference lifetime prevent mixed units and premature unloading.

R81 keeps B7 while R82 starts on B8The request retains one immutable bundle. Feature namespaces and reference lifetime prevent mixed units and premature unloading. r81 to replica: Acquire reference B7; deploy to replica: Warm and validate complete B8; replica to r81: Return pinned M7 / F4 / I7; deploy to replica: CAS active B7 to B8; r81 to store: Batch lookup schema F4; store to r81: F4 values with timestamps; r81 to replica: Score with retained M7; replica to r81: Compatible scores; policy next; r81 to replica: Release reference B7; deploy to replica: Retire B7 only after drainPARTICIPANTRequest R81PARTICIPANTServing replicaPARTICIPANTDeploymentPARTICIPANTFeature service1. Acquire reference B72. Warm and validatecomplete B83. Return pinned M7 / F4 / I74. CAS active B7 to B85. Batch lookup schema F46. F4 values with timestamps7. Score with retained M78. Compatible scores; policynext9. Release reference B710. Retire B7 only after drainsynccontrolreturn
Read each connection in order
  1. syncAcquire reference B7Request R81 → Serving replica
  2. controlWarm and validate complete B8Deployment → Serving replica
  3. returnReturn pinned M7 / F4 / I7Serving replica → Request R81
  4. controlCAS active B7 to B8Deployment → Serving replica
  5. syncBatch lookup schema F4Request R81 → Feature service
  6. returnF4 values with timestampsFeature service → Request R81
  7. syncScore with retained M7Request R81 → Serving replica
  8. returnCompatible scores; policy nextServing replica → Request R81
  9. syncRelease reference B7Request R81 → Serving replica
  10. controlRetire B7 only after drainDeployment → Serving replica

14Failure and recovery

Failure and recovery table

Timeline What user U27 sees Durable state and recovery
X9 commits, response is lost Retry may retrieve its stored session page No visibility is inferred from X9 alone; final policy rechecks still run
E55 commits, collector loses its reply Client retries E55 Durable log plus sink deduplication preserves one effect
Ranker stalls beyond 60 ms Contextual eligible fallback or shorter page Cancel expensive work; record fallback, not a fictitious model score
Policy authority is partitioned Known-safe permitted fallback or unavailable response Never promote a stale item cache into permission authority
Training crashes halfway through B8 Requests continue on B7 Incomplete artifacts remain unready; retry build before activation

Bound fallback work

A feature-store outage should not trigger 5,000 clients/s to issue unbounded independent database scans. Cap fallback lookups, use already maintained contextual pools and shed excess before ranker work begins. Temporarily stop calling a repeatedly failing retriever so it cannot consume each request’s deadline. A queue backlog delays feature updates and training. Keep its storage separate from serving memory and alert on the oldest event’s age.

Roll back the complete bundle

During a bad rollout, stop new B8 assignments and activate the previous complete compatible bundle on ready replicas. In-flight B8 requests may finish if their outputs are still permitted; an emergency policy takedown is enforced by final policy independently of the model rollback. If the old feature namespace was deleted, rollback is not just a pointer change—rebuild or serve the baseline until dependencies are ready. Test this failure, not only the happy-path activation.

15Operations, security, and cost

Assignment versus exposure

Partition experimental assignment by a stable user key, persist the assignment/version and record actual eligible exposure separately. Use user-level outcome aggregation when events from one user are correlated. Inspect new-user, language, region and other product-relevant cohorts. A treatment with 4% more clicks, twice the immediate dismissals and p99 increasing from 180 to 320 ms needs a guardrail review before launch. Its p99 still meets the stated 400-ms target, but dismissals or latency regression may violate separately agreed experiment limits. Set those limits before the experiment; one improved metric does not establish overall benefit.

Serving and feedback metrics

Monitor end-to-end/stage p95 and p99, empty-result rate, fallback fraction, candidate-source mix, feature missingness/age, bundle mismatch rejection, index freshness and feedback acknowledgement/deduplication. Record the versions and decision inputs needed to explain why I11 appeared, without logging unnecessary sensitive attributes. Protect catalog administration and model publication with separate permissions; a forged bundle or mass bot feedback is a control/data integrity threat. Apply tenant isolation if this platform serves multiple products.

Retrieval, scoring and training cost

The main cost drivers are scoring CPU, vector/index RAM, network fanout for features, event retention and repeated training scans. At our assumed 50 busy ranking cores, doubling candidates to 400 approximately doubles that scoring work before efficiency changes. Compare the incremental utility with the added 50 busy cores and tail latency. A 5.12 GB raw index is not a 5.12 GB production footprint: benchmark index expansion, replicas and overlapping rollout versions. Avoid price claims without a specific deployment quote.

Controlled rollout and recovery tests

Roll out one source/model change at a time behind a canary, replay representative known inputs, and test feature corruption, stale catalog entries, consent-off paths and log replay. Keep an independent contextual baseline deployable. Restore tests must recover version manifests and event provenance along with item vectors; a vector file without its schema cannot be trusted merely because its bytes survived.

Unbiased causal comparison

16Decision ledger and limitations

Decision table

Choice Benefit Cost / remaining risk Change trigger
Bounded retrieval before rich ranking Makes ten-million-item serving feasible Missed candidates cannot be recovered by the ranker Low candidate recall motivates another source or larger shortlist
Contextual/popularity fallback Keeps useful service during partial failure Less personal relevance and possible source bias Fall back to error if current eligibility cannot be established
Namespaced features and immutable bundles Prevents mixed units and incomplete rollout Temporary duplicate memory/storage and lifecycle work Deprecate old versions only after requests and rollback window drain
Durable validated feedback Replayable learning and measurement Hundreds of GB/day plus deduplication and deletion work Sampling/compression with measured bias and explicit guarantees
Short-lived ranked sessions Stable pagination without reranking each page Session storage and temporarily stale ordering Recompute when context changes or token expires

Model complexity and operational limits

A deep model may improve predictions, but can make results harder to explain, responses slower and rollback more involved. A simple weighted score may be the best first measured product. Similarly, caching full responses saves more CPU than caching candidates but increases stale ordering and privacy-key risks; cache scope must include user/consent context, and final policy checks still apply. A CDN key containing only the URL would be an unacceptable cross-user leak for a personalized response.

Remaining quality limitations

Remaining limitations include biased observed feedback, changing interests, adversarial content and uneven catalog coverage. Our design provides instrumentation and experiments to investigate them; it does not claim that scaling infrastructure solves relevance. If the interviewer asks for subsecond adaptation to every watch, reassess streaming feature freshness, ordering and cost while preserving event identity and feature-version semantics.

17Interview closing

Rehearse the architecture and contract

“I first build indexed contextual popularity, trustworthy response/visibility logging and a hard final policy check. At the assumed 5,000 peak requests per second, scoring the entire catalog is infeasible, so I retrieve a bounded union from several sources and richly score two hundred candidates. I batch feature reads, keep a 200-millisecond serving budget and use a contextual fallback when optional ranking components fail.

Defend the critical boundary

“The authoritative catalog and consent service decide eligibility. The model only orders allowed options. Each request pins a complete compatible model-and-feature bundle so a rollout cannot mix incompatible model expectations and feature representations. The learning path validates and deduplicates visible events, creates point-in-time examples and releases tested immutable bundles. Its queues and training jobs are outside the request path.

State the cost and next measurement

“I accept candidate-recall loss, feature freshness lag and temporary duplicate bundle memory. I measure useful viewing and satisfaction with diversity, safety and latency guardrails rather than clicks alone. My next investigation is whether each added candidate source or larger ranker improves those outcomes enough to justify its compute and operational cost.”

Answer the follow-up

Interviewer: “Personalization must stop immediately after opt-out.” Candidate: “Each new request checks current consent and discards or ignores personalized session state after opt-out. Contextual retrieval remains available. I also stop building eligible personal features and apply the defined deletion policy to retained datasets; merely hiding personalized results in the UI would leave the learning system using the same data. Previously delivered screens and in-flight authorized requests need an explicit product policy.”

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

Why not score every catalog item for every request?

Reveal a model answer

That multiplies expensive model work by the full catalog size. I retrieve a bounded set with cheaper indexes and multiple candidate sources, then apply richer scoring and final policy/diversity rules. Each stage has an explicit recall, latency, and quality tradeoff.

What the answer must demonstrate: Ranking quality is bounded by candidate recall.

Applied · Question 2

I12 scores 0.825. Why might I13 at 0.48 still be returned before it?

Reveal a model answer

I12 is already watched under our exclusion policy, so it is ineligible regardless of model score. I13 may remain a valid candidate and add diversity. Scoring orders acceptable options; it does not grant permission to violate a delivery rule.

What the answer must demonstrate: Separate eligibility from optimization.

Applied · Question 3

The online feature is minutes watched, but training used seconds. What breaks?

Reveal a model answer

The same named feature now has different meaning and scale, so the model’s learned relationship may be applied incorrectly. I version feature schema/transformations, validate compatible model bundles, and compare served feature distributions with training expectations before rollout.

What the answer must demonstrate: Check both a feature’s unit and whether its value was available when the recommendation was made.

Foundation · Question 4

A new user has no interaction history. What recommendation baseline can serve useful results?

Reveal a model answer

Use eligible contextual popularity, language/region, optional declared interests, and a diverse baseline. I avoid pretending sparse data supports confident personalization. As consented interactions accumulate, personalized retrieval can become one source rather than replacing all exploration immediately.

What the answer must demonstrate: Cold start affects users and items differently.

Follow-up · Question 5

Why can maximizing click-through rate create a misleading improvement?

Reveal a model answer

Clicks depend on what was exposed and where it appeared, and may reward curiosity or low-quality content rather than lasting utility. I evaluate causal experiment outcomes with satisfaction, quality, diversity, and operational guardrails, then inspect important cohorts rather than trust one aggregate ratio. I keep randomized-assignment analysis for the predefined population; analyzing only users who successfully saw treatment can introduce selection bias.

What the answer must demonstrate: The recommender influences the data it later learns from.

Follow-up · Question 6

The ranker has not responded after its 60-ms budget. What returns?

Reveal a model answer

I use a bounded predefined fallback over candidates that still pass current eligibility checks, such as a cached compatible score or contextual order. I record fallback/version context and preserve time for filtering and the response rather than waiting past the whole deadline.

What the answer must demonstrate: A fallback must still enforce current permissions, deletion and consent checks.

Applied · Question 7

R81 returned twenty items, but user U27 closed the app before rendering. What enters the training set?

Reveal a model answer

X9 records a returned page, not twenty visible impressions. Without a validated visibility event, I do not label these twenty items as ignored by user U27. I keep separate response, visibility and outcome records, join by authenticated page/item identity, and monitor missing telemetry so loss is not mistaken for dislike.

What the answer must demonstrate: Distinguish intent from observation and show a durable deduplication boundary.

Follow-up · Question 8

B8 activates halfway through R81. Can its old model safely use the new feature store?

Reveal a model answer

R81 keeps B7’s model M7 and feature namespace F4 throughout the request. New requests may use B8/F5; keep the old artifacts until their requests finish. If F4 is missing or incompatible, use the declared fallback rather than substitute F5. Query/item embedding encoders and the index must also be compatible; equal vector dimensions alone do not establish that.

What the answer must demonstrate: Version compatibility, lifecycle and temporal availability are separate requirements.

Blank-page exercise · 45 minutes

Build the answer yourself

Build user U27’s recommendations from one popularity query. Explain the 5,000-request/s peak, twenty result slots, feature-unit rollout race, missing visibility event, and policy-service outage.

  • State functional actions, utility/latency targets, exclusions and policy invariants.
  • Calculate full-catalog versus 200-candidate CPU and event retention.
  • Draw the baseline and identify a measured bottleneck and a misleading observation.
  • Estimate the resource cost of the four architecture changes, then trace when serving and feedback requests are acknowledged.
  • Prove bundle pinning and event deduplication under concurrent rollout/replay.
  • Close with a defensible limitation, experiment and consent-change adaptation.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Design a recommendation platformWhy have retrieval before ranking?Recall first, then reveal

Retrieve a manageable candidate set from a large catalog, then spend more compute scoring only those candidates.

Find hundreds; carefully order twenty.

Return to lesson
Design a recommendation platformWhat is training-serving skew?Recall first, then reveal

Features or their timing/meaning differ between training examples and live requests.

Same feature name must mean the same fact.

Return to lesson
Design a recommendation platformWhy are observed clicks biased?Recall first, then reveal

Users can click only items that were shown, and position/exposure affects their choices.

Exposure shapes feedback.

Return to lesson

Final revision

Summary and interview notes

Find a bounded set of candidates, check eligibility and rank them within a deadline. Keep model, features and indexes compatible throughout each request. Record what users actually saw and did, then use that evidence in a separate learning pipeline.

Remember these points

  • Rich ranking cannot recover candidates that retrieval omitted; complementary sources and controlled discovery address coverage.
  • Bundle compatibility includes transforms, feature units, defaults and query/item embedding spaces.
  • Final policy binds current consent and the exact displayed item revision; a score never grants access.
  • Returned pages, visible impressions and outcomes are different events, with scoped deduplication and current learning consent.
  • Randomized assignment supplies the main causal comparison; conditioning only on observed treatment exposure can bias it.

Interview tips

  • Compute full-catalog ranking cost, then justify the 200-candidate budget with recall and latency evidence.
  • Explain one score, one hard exclusion and one diversity decision before naming a complex model.
  • Test a missing screen-visibility event, delayed feedback received after opt-out, and a feature calculated after the recommendation from events that occurred earlier.

Important qualifications

  • Historical YouTube research illustrates two-stage design and is not evidence of today's production implementation.
  • A complete versioned model bundle prevents compatibility failures but does not prove improved satisfaction.
  • Personalization opt-out covers feature/data use under the product policy, not merely hiding a personalized UI.

Technical references

Practice marks stay in this browser.