System designby Learnastra

System-design interview · Core interviews

Design a photo-sharing service

By Anup Rai

Design durable image upload and processing, galleries, title search and follower feeds; publish only complete image variants, copy ordinary authors' photo references to followers, and merge very popular authors' photos during reads while enforcing access checks.

You will learn to

  • Separate durable photo bytes from searchable metadata and feed entries.
  • Trace one upload through processing and into a follower’s visible feed.
  • Compare write-time distribution with read-time merging using actual fanout work.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Caching: cache hits, misses, write policies and invalidation · Database indexes: B-trees, composite keys and query access · Message queues, event logs, delivery guarantees, and backpressure · Data partitioning and sharding

Workload and timing examples are interview assumptions.

01Problem and scope

A photo-sharing service stores images, publishes metadata, and lets users discover eligible photos through author galleries and a home feed. An author gallery is one author’s ordered photo list; a home feed combines photos from followed authors. Feed entries hold photo references, not duplicate image bytes. For example, publishing photo p900 produces a preview for a twenty-item feed and a larger variant for the photo page. The design must make the required variants durable before the photo becomes visible.

Scope the interview to photos, follows, title search, profile galleries, private accounts and a useful feed. Begin with recent eligible photos and allow bounded ranking over that candidate set. Upload acceptance means the original is durable; publication waits for the required image variants. This distinction prevents asynchronous processing from exposing broken feed items.

Exclude comments, tagging people, tag search, follow recommendations and cross-platform sharing from the initial design. Those are distinct products, not boxes to add without estimating them. Use hypothetical photo and user counts to compare sharding alternatives and feed strategies; the resulting design is an interview exercise. The interview will focus on safe publication, affordable delivery and controlling the work created by very popular authors.

02Functional requirements

Publishing a photo and preparing followers' feeds are separate actions. Publication makes the required images available; fanout distributes references to that published photo into followers' stored candidate lists. A feed reference helps find the photo, but the read path must still decide whether the viewer may access it.

  1. Upload original: Accepted bytes survive the agreed storage failure; UI shows processing.
  2. Publish: Required preview and large variant exist before a READY photo enters feeds.
  3. Follow: Relationship is durable; eligible new photos eventually enter the feed.
  4. Read a feed page: Up to 20 eligible unique photo IDs with stable pagination semantics.
  5. Search title: Results may lag, but deletion/privacy checks happen before disclosure.
  6. Delete: Metadata reads stop exposing the photo after the database commits deletion; existing media links have a stated short lifetime.

User actions and private accounts

The author can request an upload, resume or retry it, observe processing, and obtain a published photo page. The viewer can follow or unfollow an author, page through a feed, search visible titles and view author galleries. Owners may delete photos. Private accounts require approved follower membership before metadata or media access is granted; an old feed reference is not itself an authorization grant.

Retry, invalid-image and feed behavior

An invalid image is rejected with a reason rather than kept processing forever. A retry of the same upload request returns the same photo identity. Repeated fanout events do not duplicate a photo in the viewer's inbox. Unfollow hides that author's candidates on subsequent reads even if asynchronous cleanup has not removed the references. Exact like counts and personalized machine-learning ranking are extensions; the first ranking uses recency with optional bounded relevance features.

03Non-functional requirements

Feed metadata and image bytes follow different request paths, so they have separate latency and access guarantees. A media-delivery token is a short-lived credential for requesting a permitted image variant. Allowing delivery under that credential until expiry reduces repeated permission-service calls, but creates the revocation window stated below.

  1. Feed latency: Feed metadata p95 below 200 ms inside the serving region.
  2. Availability: 99.9% eligible feed success inside the serving region.
  3. Media latency: Preview time to first byte below 300 ms from an available nearby delivery cache; measure it separately from feed metadata.
  4. Processing time: 95% of valid ordinary images become READY within 30 seconds under planned peak load.
  5. Feed freshness: Ordinary-author propagation should usually finish within five seconds. Expose backlog; a slightly older authorized feed is acceptable.
  6. Durability: Acknowledged original storage and metadata commits survive one node or availability-zone failure through correctly placed/configured replicas.
  7. Regional recovery: Initially use asynchronous replication with an explicit measured recovery point and a one-hour restore target. Derivatives and feeds are rebuildable; lost originals cannot be recovered from feed IDs.
  8. Private-media revocation: Delivery tokens last at most 60 seconds. An already issued token may remain usable for that interval; immediate revocation requires current authorization checks on every edge request.

Publication and authorization rules

Rule Required behavior
READY publication Reference verified durable variants and a retained original.
Metadata access Check current deletion state and membership.
Authority partition Reject new private grants; a cached feed is not permission.
Immediate media revocation option Pay the added edge latency and permission-service dependency for per-request checks.

The failure contract is separate from the availability percentage. Proper replication supports the stated node/zone guarantee; it does not justify a claim of “100% reliability.”

04Capacity estimates

Workload assumptions and arithmetic

Assume 500 million registered users, one million daily active users, two million photos/day and 200 KB average originals. Add ten feed opens per active user/day with twenty 50 KB previews per page. These are explicit exercise assumptions, not observed traffic. 2M / 86,400 = 23.1 uploads/s; fivefold peak is about 116/s. Feed requests average 10M / 86,400 = 116/s, peaking near 579/s.

Worked estimates

Quantity Calculation What it changes
Original bytes 2M × 200 KB = 400 GB/day Bulk storage must grow independently
Ten-year originals 400 GB × 365 × 10 = 1.46 PB Retention dominates long-term bytes
Preview delivery 10M × 20 × 50 KB = 10 TB/day Geographic caches reduce origin traffic
Mean preview egress 10 TB / 86,400 ≈ 116 MB/s Delivery is much larger than API payload
Follow edges 500M × 500 × 16 B = 4 TB raw Both directions and indexes add cost
Metadata illustration 2M × 284 B × 365 × 10 ≈ 2.07 TB Count bytes consistently before indexes

Capacity implications and limits

The 284-byte record-size estimate is illustrative; actual IDs, strings and indexes change it. Derivatives, replicated copies and backups are additional. If an ordinary author has 300 active followers, 23.1 average photo publications/s yield about 6,930 inbox inserts/s before celebrity exceptions. At peak, about 34,800/s. A single fifty-million-follower author breaks this average immediately. A content delivery network (CDN) that serves 90% of requested image bytes from its caches reduces the 10 TB/day preview origin demand toward 1 TB/day, but viewers still receive 10 TB/day and caches still incur that delivery cost.

For a minimal user-record estimate, assume 500 million users × 68 bytes = 34 GB of raw fixed fields. That is an arithmetic floor, not the size of a production user table: variable profile fields, indexes, access-control records and replicas add bytes. It reinforces why media capacity and metadata capacity need separate estimates.

05APIs and contracts

Request and response example

The author calls POST /v1/photo-uploads with request key upload-90 and {"title":"Sunrise","bytes":200000,"visibility":"followers","checksum":"H900"}. The response names photoId:p900, uploadId:up900, a generation-specific object target and an expiry. The upload capability is issued only after authenticating the owner and is limited to this object generation and operation. A presigned URL is a bearer credential, not a later proof of the uploader’s identity. Bind supported length/checksum conditions cryptographically or through the upload policy, and recheck the accepted object at completion. Completion is a separate authenticated request; possessing a storage upload token cannot publish metadata directly.

Interface contracts

API Result and error behavior
POST /photo-uploads/up900/complete Verify object; 202 with processing state; duplicate completion reuses state
GET /photos/p900/status Owner sees uploading/processing/ready/failed
PUT /following/u17 Idempotently follow the author, or create a pending request for private accounts
GET /feed?cursor=<token>&limit=20 Visible metadata, short-lived media tokens, next cursor
GET /users/u17/photos?before=<time,id> Author/time gallery page
GET /photos/search?q=sunrise&cursor=... Title-index candidates filtered by current visibility
DELETE /photos/p900 Owner-authorized tombstone; repeated delete is harmless

Validation and response semantics

The feed cursor refers to a bounded candidate snapshot and offset or stable ranking key; it is scoped to the viewer and expires. Deletions may make a page shorter, so the server may fetch extra candidates within a work limit. Invalid formats return 400, oversized files 413, request-key payload mismatch 409 and temporary capacity failures 429/503. A client never invents a new upload identity merely because a completion response timed out.

06Data model and access patterns

The model separates the authoritative photo from the work needed to publish and distribute it. A manifest is the list of accepted image variants and their storage references; readers use it to select complete outputs. The outbox stores processing or publication work in the same database transaction as the photo change, while Feed stores rebuildable per-viewer references rather than image bytes.

Record and fields Responsibility / constraint
Photo(photoId, ownerId, createdAt, title, visibility, state, sourceGeneration, originalKey, originalVersionId, manifest, version) The manifest identifies preview and large-image objects for one published generation.
Upload(uploadId, ownerId, requestKey, photoId, checksum, leaseUntil) Owns upload retry state.
Follow(followerId,authorId,status,version) Owns relationships.
Outbox(eventId,photoId,generation,type) Records durable work.
Feed(viewerId,photoId,sortKey) Derived; unique by viewer/photo identity.

The author's gallery query needs (ownerId,createdAt DESC,photoId DESC), not merely an index on photo ID. The viewer's follow list needs (followerId,authorId); publication fanout needs the reverse (authorId,followerId) access path. The title index is another derived view; it cannot authorize a private photo. A gallery query reads WHERE ownerId='u17' AND (createdAt,photoId)<(:t,:id) ORDER BY createdAt DESC,photoId DESC LIMIT 20.

Partition primary photos by a hash of photo ID when independent growth requires it. Maintain an author/time index to avoid querying every photo shard for the author's gallery. Partition inboxes by viewer, with bounded recent retention; shard huge follower lists into pages. Time buckets can reduce old-data scanning but should be combined with hashing or author keys so the newest bucket does not become the only write target. Metadata and outbox changes for a photo share one authoritative transaction. Cross-partition feed inserts are asynchronous, never part of the publication commit.

07Basic working design

Synchronous publication on one server

Start with one app and a SQL database containing photo metadata, follows and original/preview bytes for a small corpus. The author uploads, the app validates and generates the two required variants synchronously, then commits the photo and bytes together. Only after commit does p900 become visible. The upload response can be slow, but the single commit boundary is easy to understand. A failed decode produces no ready photo.

Pull-on-read feed assembly

When the viewer opens a home feed, read the viewer’s followed-author list, query recent photos from those authors, filter visibility, sort by time and return the first twenty. Fetching a bounded number of recent rows per author is a straightforward implementation. The author's gallery is one indexed query. Title search can initially use a modest database text index rather than a dedicated search cluster.

What the baseline buys and where it stops

For a tiny active population this is operationally attractive: one backup covers metadata and original bytes, one transaction publishes, and debugging p900 is local. The baseline's weaknesses are synchronous image processing, centralized media egress and repeated multi-author feed work. It is intentionally a working product rather than an unfinished drawing. Later changes must still return complete, authorized photos while reducing image transfer and repeated feed assembly in requests.

architecture · baselineLocal publication and pull-on-read feed

One database commits the image and metadata; feed work repeats across followed authors.

Local publication and pull-on-read feedOne database commits the image and metadata; feed work repeats across followed authors. client to app: 1. Upload or open feed; app to db: 2. Commit validated photo; app to db: 3. Query followed authors; app to client: 4. Return page and image bytes1. Upload or open feed2. Commit validated photo3. Query followed authors4. Return page and imagebytesACTORUploader and viewerclientsSERVICEPhoto and feedapplicationSTORESQL photos, followsand image bytessync
Read each connection in order
  1. sync1. Upload or open feedUploader and viewer clients → Photo and feed application
  2. sync2. Commit validated photoPhoto and feed application → SQL photos, follows and image bytes
  3. sync3. Query followed authorsPhoto and feed application → SQL photos, follows and image bytes
  4. sync4. Return page and image bytesPhoto and feed application → Uploader and viewer clients

08Find the baseline flaws

Bottleneck / counterexample Evidence and design consequence
Repeated feed candidate work Suppose the viewer follows 500 authors and the baseline retrieves each author's latest 100 photos before ranking. That is 50,000 candidates to produce twenty results. At 579 peak feed requests/s, repeating this policy examines roughly 29 million candidate rows/s before loading metadata for those photo IDs. Even a well-indexed table cannot erase that repeated work. The first improvement is to bound or precompute candidates, not merely add a ranking service.
Upload/read resource contention Uploads create another bottleneck. At 116 uploads/s and four seconds of transfer/processing time, roughly 464 uploads are active. Sharing a hypothetical 500-connection or worker budget with feed reads can starve the read path. The exact limit is implementation-dependent; the point is to measure occupancy and isolate work with different duration, rather than assert every server has the same universal connection limit.
Premature READY publication A correctness counterexample appears after naive asynchronous resizing. The API inserts READY metadata and pushes a job, then the worker crashes before writing the preview. The viewer sees a broken feed item. Or the database commits PROCESSING but the separate queue send fails, leaving the photo stuck forever. These are different failures. To prevent the broken photo, a worker verifies all required images before committing their manifest. To prevent lost work, store the processing event with the state change in a transactional outbox and send it after commit.

09Improve the design, step by step

  1. Move original bytes to private object storage and separate upload/read pools. The trigger is growing byte retention plus hundreds of slow uploads. Direct constrained uploads remove bulk transfer from feed servers; dedicated completion APIs verify objects. This improves read isolation and independent storage scaling. It costs two-store coordination, token management and orphan cleanup. Keeping bytes in the database remains simpler for small corpora; do not split before the operational benefit is real.

  2. Process variants asynchronously with a durable outbox. The trigger is decode latency and variable image complexity. Commit PROCESSING with an event, then workers create immutable generation-specific variants and atomically publish a complete manifest. This gives predictable upload acceptance and recoverable work. It costs queues, worker capacity, duplicate handling and a visible processing state. Synchronous processing is preferable when bounded small inputs reliably fit the response budget.

  3. Introduce hybrid feed preparation. The trigger is repeated 50,000-candidate assembly. Ordinary authors' photo IDs are inserted into active followers' inboxes; celebrity photos stay in author lists and merge on read. This makes ordinary feed reads cheap while avoiding fifty million writes for one celebrity upload. It costs inbox storage, fanout checkpoints and two-path deduplication. Pure pull is attractive for inactive viewers or small follow lists; pure push works when follower counts and write amplification remain bounded.

  4. Add media delivery caches, metadata partitions and replicated authority. The triggers are roughly 10 TB/day preview delivery, retained metadata growth and zone-failure durability. Immutable variants are cached near viewers; logical partitions distribute records and queries; replicated leaders protect acknowledged changes. Costs include origin-fill bursts, routing epochs, duplicate storage and authorization-token lifetime. A single larger replicated database remains viable until measurements justify partitions; SQL is not ruled out by the product's name.

Each change preserves the invariant that only a committed READY manifest enters candidate feeds. None permits a fanout worker or CDN to decide independently that private content is public.

10Detailed architecture

Upload control and metadata authority

The edge routes upload-control traffic to an upload API and feed/search traffic to a read API. The phone sends bytes directly to private object storage using a constrained upload target. The metadata database stores uploads, photos, follows and publication events. Each partition’s replicated leader commits changes in order. The diagram shows representative groups, not one global lock for every photo.

Processing and feed workers

An outbox relay forwards processing and publication events to a durable work queue. Image workers write immutable variants, then ask the database partition storing the photo to publish its manifest. Feed workers consume publication events and page through follower lists into viewer inboxes. Search indexing and author-list updates consume the same committed publication stream, with idempotent event identity. They may lag without changing p900's authoritative state.

Authorized feed and media reads

Read APIs combine inbox candidates with celebrity author lists, batch-load metadata, check current visibility/membership, and return a bounded page. A separate media edge validates short-lived access tokens and serves CDN-cached bytes or fetches the private origin. Thus authorization is in front of delivery, not an optional caption beside a public bucket. Ranking may degrade to authorized recency order; permission checks may not degrade to “allow all.” Logical partition routing is versioned, and migrations fence old owners before accepting writes at new locations.

architecture · finalPublish once, prepare candidates, authorize delivery

Only a committed READY manifest emits the publication event. Inbox and search entries are candidates; the read path still authorizes the photo.

Publish once, prepare candidates, authorize deliveryOnly a committed READY manifest emits the publication event. Inbox and search entries are candidates; the read path still authorizes the photo. client to edge: 1. Upload control / page request; edge to upload: 2a. Admit upload session; edge to read: 2b. Request visible candidates; upload to meta: 3. Reserve / pin verified source version; client to objects: 4. Upload scoped original g1; meta to replicas: 5. Replicate authoritative state; meta to queue: 6. Relay committed outbox; queue to image: 7. Process current attempt; image to objects: 8. Write immutable variants; image to meta: 9. Guarded READY + outbox; queue to fanout: 10. Publish references and indexes; fanout to views: 11. Idempotent candidate updates; read to views: 12. Merge bounded candidates; read to meta: 13. Load photo metadata and check access; read to client: 14. Page and scoped media token; client to delivery: 15. Present token + bound session; delivery to objects: 16. Cache miss: private origin1. Upload control / pagerequest2a. Admit upload session2b. Request visible candidates3. Reserve / pin verified sourceversion4. Upload scoped original g15. Replicate authoritative state6. Relay committed outbox7. Process current attempt8. Write immutable variants9. Guarded READY + outbox10. Publish references andindexes11. Idempotent candidateupdates12. Merge bounded candidates13. Load photo metadata andcheck access14. Page and scoped mediatoken15. Present token + boundsession16. Cache miss: private originACTORMobile and webclientsSERVICEEdge request routingG1SERVICEUpload control APIG1SERVICEFeed, gallery andsearch APIG1STOREPartitioned metadataauthorityG2STOREMetadata replicasG2STOREPrivate original andvariant storeG2QUEUEOutbox relay andwork queueG3WORKERImage processingworkersG3WORKERFeed and indexworkersG3STOREInbox, author andtitle indexesG3CACHEAuthorized mediaedge and CDNG4syncreplicationasyncG1 Request and identity boundaryG2 Authoritative metadata and mediaG3 Asynchronous processing and viewsG4 Media authorization and caching
Read each connection in order
  1. sync1. Upload control / page requestMobile and web clients → Edge request routing
  2. sync2a. Admit upload sessionEdge request routing → Upload control API
  3. sync2b. Request visible candidatesEdge request routing → Feed, gallery and search API
  4. sync3. Reserve / pin verified source versionUpload control API → Partitioned metadata authority
  5. sync4. Upload scoped original g1Mobile and web clients → Private original and variant store
  6. replication5. Replicate authoritative statePartitioned metadata authority → Metadata replicas
  7. async6. Relay committed outboxPartitioned metadata authority → Outbox relay and work queue
  8. async7. Process current attemptOutbox relay and work queue → Image processing workers
  9. sync8. Write immutable variantsImage processing workers → Private original and variant store
  10. sync9. Guarded READY + outboxImage processing workers → Partitioned metadata authority
  11. async10. Publish references and indexesOutbox relay and work queue → Feed and index workers
  12. async11. Idempotent candidate updatesFeed and index workers → Inbox, author and title indexes
  13. sync12. Merge bounded candidatesFeed, gallery and search API → Inbox, author and title indexes
  14. sync13. Load photo metadata and check accessFeed, gallery and search API → Partitioned metadata authority
  15. sync14. Page and scoped media tokenFeed, gallery and search API → Mobile and web clients
  16. sync15. Present token + bound sessionMobile and web clients → Authorized media edge and CDN
  17. sync16. Cache miss: private originAuthorized media edge and CDN → Private original and variant store

11Write path and acknowledgement

The upload protocol publishes only a verified image manifest and records downstream work durably. Photo p900, upload up900 and request upload-90 provide concrete identifiers for the transitions.

  1. The author authenticates as u17 and creates upload up900 using upload-90. A transaction reserves photo p900 in UPLOADING with a checksum, generation and deadline.
  2. The upload client uploads originals/p900/g1. The storage token cannot write other users' keys. A failed transfer resumes/retries the same session rather than creating an unrelated photo.
  3. Completion verifies expected size/checksum and permitted image format, then commits PROCESSING and outbox event process-p900-g1. Return 202: original accepted, publication pending.
  4. A relay publishes the event. Worker W1 decodes with pixel/dimension limits, removes disallowed metadata such as location data when policy requires it, and writes preview and large variants under an attempt-specific immutable prefix.
  5. W1 verifies every required output and submits a manifest plus its claimed job generation. The metadata owner atomically changes PROCESSING to READY only for the current valid generation and inserts ready-p900-g1 in the outbox.
  6. Feed workers read this committed event, page through ordinary active followers, and insert (viewer31,p900) if absent. Title search and author/time indexes receive idempotent updates.
  7. The author's status request returns READY. If completion or publication responses were lost, retry reads the existing session/photo state. Nothing in the protocol requires generating a second p900.

Unused attempt outputs never appear in the manifest. Delete them only after the metadata state prevents any current worker from publishing them.

A generation-shaped key is not automatically immutable in an object store. S3 presigned upload URLs can be reused before expiration and can replace the current object at the key. One safe implementation uses a versioned bucket, verifies an exact accepted VersionId and checksum at completion, stores that VersionId in the photo record, and makes every worker read that exact version. Later uploads to the same key cannot change the source already accepted. Alternatively enforce a conditional create-only upload with the required signed checksum. Retain the accepted version through lifecycle rules; default current-key GETs and blanket noncurrent-version expiration would break the version-pinned design.

12Read and delivery path

Feed assembly chooses candidate IDs, checks current access, then issues bounded media grants. A twenty-item request illustrates these responsibilities without treating an old inbox entry as permission.

  1. The viewer authenticates and requests a twenty-item feed page. The read API validates the cursor's viewer identity and snapshot lifetime.
  2. It loads a bounded inbox window and recent photos from followed high-fanout authors. It merges by sort key, deduplicates photo IDs and applies a candidate limit to keep one request's work bounded.
  3. Batch-fetch photo metadata by ID, using caches for immutable fields while rechecking authoritative deletion/visibility and current membership according to the private-access contract. Remove p900 if the author deleted it or the viewer no longer has access.
  4. Rank eligible candidates by recency and bounded relevance signals. Return twenty results or a shorter page with a cursor if the bounded candidate window contains too few eligible items. Do not issue unbounded fanout queries just to fill every page perfectly.
  5. For p900, issue a media token scoped to the viewer or the authorized session, object generation, variant and 60-second expiry. Return the title, owner, dimensions and preview route.
  6. The viewer's device requests the media edge. It validates the token before serving cached preview bytes; a miss fetches the private origin. The original remains inaccessible unless separately authorized.

Polling, long polling or push notifications can tell the viewer that new items are available. These delivery mechanisms are separate from database fanout-on-write. Coalesce notifications for busy followers rather than pushing one user-interface refresh for every photo. Gallery and title-search reads use their own indexes but finish with the same visibility checks.

13Correctness deep dive

Retries need a publication guard

A worker can finish writing image variants, then lose its queue acknowledgment. The queue may therefore send the job again. “At least once” means a job can be delivered again after a timeout; it does not mean the photo should be published twice. Use a job generation and a lease token whose validity is checked by the metadata authority during publication.

Operation Authority check Durable effect
Claim processing Photo PROCESSING, no active valid claim Save attempt token 41 and lease
Reclaim after expiry Token 41 expired, still PROCESSING Save token 42; old token becomes invalid
Publish manifest Current token matches, lease valid, photo not deleted READY plus one unique publication outbox event
Repeat publication Same committed generation already READY Return existing manifest, no second event
Insert feed reference Unique (viewerId,photoId) absent One candidate reference; duplicate is harmless

Stale-worker publication proof

Fanout checkpoint and deduplication proof

sequence · stale-encoderA stale image worker cannot publish or overwrite

Attempt-specific objects prevent stale byte writes; the metadata token comparison prevents stale publication.

A stale image worker cannot publish or overwriteAttempt-specific objects prevent stale byte writes; the metadata token comparison prevents stale publication. w1 to meta: Claim token 41; w1 to obj: Write attempt-41 preview; w2 to meta: After expiry: claim token 42; w2 to obj: Write complete attempt-42 set; w2 to meta: Publish manifest under token 42; meta to w2: READY plus publication outbox; w1 to meta: Late publish under token 41; meta to w1: Reject obsolete tokenPARTICIPANTWorker W1PARTICIPANTPhoto authorityPARTICIPANTWorker W2PARTICIPANTObject store1. Claim token 412. Write attempt-41 preview3. After expiry: claim token424. Write complete attempt-42set5. Publish manifest undertoken 426. READY plus publicationoutbox7. Late publish under token418. Reject obsolete tokensyncreturn
Read each connection in order
  1. syncClaim token 41Worker W1 → Photo authority
  2. syncWrite attempt-41 previewWorker W1 → Object store
  3. syncAfter expiry: claim token 42Worker W2 → Photo authority
  4. syncWrite complete attempt-42 setWorker W2 → Object store
  5. syncPublish manifest under token 42Worker W2 → Photo authority
  6. returnREADY plus publication outboxPhoto authority → Worker W2
  7. syncLate publish under token 41Worker W1 → Photo authority
  8. returnReject obsolete tokenPhoto authority → Worker W1

14Failure and recovery

Failure / trigger User outcome, surviving state and recovery
Image worker crash Original g1 and PROCESSING state survive. Another worker claims a new attempt after the lease expires, regenerates variants and tries the guarded publication transaction. The author sees processing longer; the viewer sees no broken READY entry. If decoding repeatedly fails, mark a durable failed state and expose a useful error rather than retrying forever.
Metadata zone failure or partition A surviving majority can elect a leader and preserve committed READY manifests. A minority cannot publish or grant new private access. Existing short-lived media tokens remain valid until their declared expiry; after that the edge cannot mint replacements without authorization. A full-region outage has the separately stated recovery window. CDN copies do not replace backups of originals or ownership metadata.
Celebrity burst One fifty-million-follower post must not enqueue fifty million urgent writes onto the ordinary path. Classification sends it to the author-list path; read caches share popular metadata and preview bytes. If ordinary fanout backlog grows, prioritize active viewers and maintain a bounded catch-up window. Return an older authorized feed with a freshness indicator while publication progresses.
Ranking or search outage Feed reads fall back to authorized recency candidates; title search may return a temporary error rather than leak unfiltered cached results. Deletion first tombstones metadata, then asynchronously removes indexes and media. Old inbox/search entries are harmless references only because read-time policy is enforced. Already downloaded images remain beyond the service's revocation control.

15Operations, security, and cost

Processing, delivery and feed signals

Track upload-to-ready percentiles, oldest processing lease, invalid-image rate, whether every manifest references existing variants with matching checksums, and whether backups restore the original images successfully. Feed metrics include p95/p99 metadata latency, candidate count, per-author fanout work, publication lag and fraction of requests using celebrity merges. Delivery metrics include byte-hit ratio, origin bandwidth, token failures and denied private requests. A simple API success counter would miss most of these user-visible failures.

Storage, egress and inbox cost

Cost depends on how long originals are retained, the number and sizes of derived image variants, bytes delivered to viewers, and the number of follower-inbox references written per photo. If each candidate reference occupies an illustrative 32 bytes, 300 follower references cost 300 × 32 = 9.6 KB per ordinary photo, versus a 200 KB original. Fifty million references cost 1.6 GB for one celebrity photo before indexes and replicas. That arithmetic explains a hybrid policy better than an arbitrary celebrity label. Determine the push/pull threshold from expected active follower reads during the useful feed window and measured merge cost.

Untrusted image handling

Exercise malformed image headers, decompression bombs, oversized dimensions, expired upload tokens and forbidden object paths. Remove unnecessary location metadata according to product policy. Do not log private media tokens. Enforce object-store access policies and account ownership at completion, not only when upload starts.

Migration and failure drills

For a partition migration, copy records and author indexes, replay changes, compare sample queries, fence the old epoch and cut over routing. Test W1/W2 lease races, a repeated fanout page, deleted photos in cached feeds, and restoration of originals with manifests. Roll out ranking separately from correctness-sensitive visibility filtering so a model change cannot bypass access checks.

16Decision ledger and limitations

Partitioning chooses which storage group owns a photo or an index entry. Feed preparation chooses when references are copied or merged for viewers. These are independent decisions: balancing primary photo records does not by itself make an author's gallery query local or bound a celebrity's fanout work.

Choice Benefit Cost and consequence Change trigger
Hash photo-ID primary records Spreads different photos and bytes Gallery requires author/time index Owner-local transactions dominate access
Owner-based partition Gallery locality Prolific/hot authors skew one owner Split large owners across time buckets
Hybrid feed candidates Cheap ordinary reads, bounded celebrity writes Two paths, checkpoints and deduplication Workload shifts toward mostly inactive viewers
Immutable variant manifests Safe retry and cache identity Orphan attempts and retained originals Strongly transactional media storage simplifies it
Short-lived private media tokens Delivery edge avoids central check per byte request Revocation bounded by token lifetime Immediate revocation becomes mandatory

Time-sortable photo IDs can include timestamp, generator identity and per-tick sequence. Each generator must handle clock rollback and sequence exhaustion without issuing duplicates. An ID format with a 31-bit seconds field has a finite time horizon, and its 9-bit sequence permits only 512 IDs per second within one allocation scope, so average 23/s does not justify safety under peaks or multiple generators. Use a proven larger scheme or allocated IDs with explicit authority.

Disjoint odd/even database sequences are another allocation alternative. Their ranges must remain disjoint through failover; standby promotion cannot reset a sequence and reuse values. A logical partition map is more flexible than hard-coded id % currentServerCount, but moving it requires a fenced migration, not just editing a configuration file. LRU metadata caching is reasonable when measured locality supports it; popularity and byte size may justify admission limits beyond recency alone.

17Interview closing

“I designed photos, follows, galleries, title search and a feed with durable publication and authorized delivery. The assumptions give about twenty-three average uploads per second but 400 GB of new originals and ten terabytes of preview delivery per day. I start with a working transactional version, then move bytes and resizing out of feed servers, add a durable outbox, and prepare ordinary followers' candidate lists.

“The hard guarantee is that a READY photo references a complete verified manifest, and a stale worker cannot replace the current generation. Feed delivery is eventually updated and idempotent; it never grants access by itself. I use hybrid fanout because a fifty-million-follower author makes per-follower writes unreasonable. Media caches serve immutable variants, while authorization happens before delivery and private tokens have an explicit sixty-second lifetime.

“I accept bounded feed staleness and some extra index complexity. The next measurements are candidate-merge cost, fanout backlog for active users, and origin bytes after cache loss.”

If the interviewer changes the feed to a highly personalized ranking, keep candidate generation and current visibility as separate stages. Add a versioned ranking model over bounded eligible candidates, evaluate relevance and latency, and stabilize pagination. Do not let a ranking score become evidence of permission or replace the publication invariant.

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

What are the three different things you store for one photograph?

Reveal a model answer

“The original bytes, metadata describing ownership and state, and references in author or follower lists. Derivatives are rebuildable versions of the original; feed entries are candidate references. Losing each has a different recovery story.”

What the answer must demonstrate: Do not store full images in every follower feed.

Applied · Question 2

A viewer follows 500 authors. How would you assemble a twenty-item photo feed, and when would you precompute it?

Reveal a model answer

“Initially I query recent author lists and merge a bounded set. If repeated reads make that too expensive, ordinary authors distribute IDs into active followers’ inboxes. I merge celebrity-author lists on read and filter current permissions before ranking.”

What the answer must demonstrate: Explain fanout’s unit of work before choosing it.

Applied · Question 3

The resize worker crashes after producing one image variant. What happens?

Reveal a model answer

“The photo remains PROCESSING, with its original retained durably. The replacement reads the exact verified source version recorded at completion but writes to its own attempt-specific output paths. It verifies every required variant, then atomically checks its current worker token before publishing READY plus the feed outbox event. The old attempt cannot overwrite the accepted paths or win publication after its token is replaced.”

What the answer must demonstrate: An output object alone must not imply ready metadata.

Foundation · Question 4

Why does a photo service need an author/time index in addition to a PhotoID primary key?

Reveal a model answer

“A PhotoID lookup retrieves one known photo. A profile asks for an author’s newest photos, so it needs an owner/time access path, such as (ownerId, createdAt, photoId). Hashing primary records otherwise scatters that range query. I add the index because of the query shape, not because photo IDs are insufficiently unique.”

What the answer must demonstrate: Ordering and locating are different tasks.

Follow-up · Question 5

A private photo is deleted after its ID entered a viewer’s feed cache. How do you prevent the stale candidate from disclosing it?

Reveal a model answer

“An inbox entry is only a candidate. Before returning its metadata or issuing a media grant, I check current deletion and membership at the authority. Asynchronous cleanup removes stale references but is not the permission boundary. Previously issued private media tokens remain usable for up to the stated 60 seconds; immediate revocation would require current checks at the delivery edge.” The edge must also check any claimed viewer/session binding against the authenticated requester; verifying a token signature alone does not enforce that binding.

What the answer must demonstrate: Do not use one consistency slogan for every read.

Follow-up · Question 6

What does a 10 TB/day preview estimate tell you?

Reveal a model answer

“Media delivery is a separate bandwidth path. I consider derivative size and distributed caching, then measure origin byte-hit ratio. It does not mean metadata needs the same capacity or that a CDN eliminates viewer traffic.”

What the answer must demonstrate: Separate payload estimates from replicated capacity.

Applied · Question 7

A paused resize worker wakes after a replacement published. What stops it corrupting the photo?

Reveal a model answer

The metadata owner checks the worker token atomically when changing PROCESSING to READY. The replacement has a new token, so the old worker cannot publish. Crucially, each attempt writes immutable object names; otherwise the stale worker could overwrite accepted bytes even if its metadata update were rejected.

What the answer must demonstrate: Reject the old worker’s manifest update and prevent it from overwriting the accepted image objects.

Follow-up · Question 8

Show the cost that makes a hybrid feed worthwhile.

Reveal a model answer

At 300 active followers, one 32-byte reference per follower is about 9.6 KB per photo. At fifty million followers it becomes 1.6 GB before indexes and replicas. I keep that large author’s recent list and merge it on active readers’ requests, while ordinary authors benefit from prepared inboxes.

What the answer must demonstrate: Do not confuse fanout-on-write with WebSocket or push notification transport.

Blank-page exercise · 45 minutes

Build the answer yourself

Design photo uploads, galleries and a follower feed. Derive the storage and delivery workload, evolve from a single-server baseline, then handle a 50-million-follower author and a resize worker that resumes after its replacement publishes.

  • Name original, derivative, metadata, and feed reference.
  • Calculate uploads, retained originals, and preview delivery bytes.
  • Show p900 state transitions and the durable event handoff.
  • Compare fanout-on-write/read with an actual follower count.
  • Keep author/time retrieval and photo lookup distinct.
  • Explain deleted-photo behavior despite a stale inbox.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Design a photo-sharing serviceWhat does fanout-on-write copy?Recall first, then reveal

Photo references into eligible followers’ candidate lists, not complete image bytes.

Many inboxes; one original.

Return to lesson
Design a photo-sharing serviceWhy keep ownerId + time after hashing photo IDs?Recall first, then reveal

Because fetching one author’s recent photographs is a different access path from finding one photo by ID.

Identity finds one; index finds a list.

Return to lesson
Design a photo-sharing serviceA preview object exists while photo metadata is PROCESSING. May the feed expose it?Recall first, then reveal

No. The worker verifies all required outputs and commits the guarded READY manifest before emitting the publication event.

A file is not a published photo.

Return to lesson

Final revision

Summary and interview notes

A photo service publishes verified media manifests and distributes candidate references, then authorizes metadata and byte delivery separately. Hybrid feeds reduce ordinary read work without turning one popular author into tens of millions of mandatory inbox writes.

Remember these points

  • At the stated workload, originals add 400 GB/day while previews deliver 10 TB/day; storage and delivery need separate capacity plans.
  • Pin the exact verified source version so a reusable upload credential cannot change an accepted image.
  • A current worker token and attempt-specific immutable output names protect both manifest publication and external bytes.
  • Copy ordinary authors’ photo IDs into active followers’ lists. Merge celebrity lists during reads, limit candidates, and save progress only after inserts are safe to repeat.
  • A stale inbox is not permission, and viewer-bound media tokens require the edge to verify the matching identity.

Interview tips

  • Show why 500 authors × 100 photos is expensive before introducing prepared inboxes.
  • Resume an old resize worker after a replacement publishes; protect both the pointer and object names.
  • Ask whether an upload URL can be replayed and whether a media token is bearer-only or actually identity-bound.

Important qualifications

  • S3 presigned URLs are reusable bearer capabilities until their applicable expiry; a named generation alone is not immutable storage.
  • The 60-second media-token contract allows a bounded revocation delay; immediate revocation needs a different serving check.
  • A derived gallery/title index may lag; it must still filter current publication and privacy state.

Technical references

Practice marks stay in this browser.