System-design interview · Core interviews
Design a photo-sharing service
Design durable image upload and processing, galleries, title search and follower feeds; publish only complete image variants, copy ordinary authors' photo references to followers, and merge very popular authors' photos during reads while enforcing access checks.
You will learn to
- Separate durable photo bytes from searchable metadata and feed entries.
- Trace one upload through processing and into a follower’s visible feed.
- Compare write-time distribution with read-time merging using actual fanout work.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Caching: cache hits, misses, write policies and invalidation · Database indexes: B-trees, composite keys and query access · Message queues, event logs, delivery guarantees, and backpressure · Data partitioning and sharding
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Problem and scope
A photo-sharing service stores images, publishes metadata, and lets users discover eligible photos through author galleries and a home feed. An author gallery is one author’s ordered photo list; a home feed combines photos from followed authors. Feed entries hold photo references, not duplicate image bytes. For example, publishing photo p900 produces a preview for a twenty-item feed and a larger variant for the photo page. The design must make the required variants durable before the photo becomes visible.
Scope the interview to photos, follows, title search, profile galleries, private accounts and a useful feed. Begin with recent eligible photos and allow bounded ranking over that candidate set. Upload acceptance means the original is durable; publication waits for the required image variants. This distinction prevents asynchronous processing from exposing broken feed items.
Exclude comments, tagging people, tag search, follow recommendations and cross-platform sharing from the initial design. Those are distinct products, not boxes to add without estimating them. Use hypothetical photo and user counts to compare sharding alternatives and feed strategies; the resulting design is an interview exercise. The interview will focus on safe publication, affordable delivery and controlling the work created by very popular authors.
02Functional requirements
Publishing a photo and preparing followers' feeds are separate actions. Publication makes the required images available; fanout distributes references to that published photo into followers' stored candidate lists. A feed reference helps find the photo, but the read path must still decide whether the viewer may access it.
- Upload original: Accepted bytes survive the agreed storage failure; UI shows processing.
- Publish: Required preview and large variant exist before a READY photo enters feeds.
- Follow: Relationship is durable; eligible new photos eventually enter the feed.
- Read a feed page: Up to 20 eligible unique photo IDs with stable pagination semantics.
- Search title: Results may lag, but deletion/privacy checks happen before disclosure.
- Delete: Metadata reads stop exposing the photo after the database commits deletion; existing media links have a stated short lifetime.
User actions and private accounts
The author can request an upload, resume or retry it, observe processing, and obtain a published photo page. The viewer can follow or unfollow an author, page through a feed, search visible titles and view author galleries. Owners may delete photos. Private accounts require approved follower membership before metadata or media access is granted; an old feed reference is not itself an authorization grant.
Retry, invalid-image and feed behavior
An invalid image is rejected with a reason rather than kept processing forever. A retry of the same upload request returns the same photo identity. Repeated fanout events do not duplicate a photo in the viewer's inbox. Unfollow hides that author's candidates on subsequent reads even if asynchronous cleanup has not removed the references. Exact like counts and personalized machine-learning ranking are extensions; the first ranking uses recency with optional bounded relevance features.
03Non-functional requirements
Feed metadata and image bytes follow different request paths, so they have separate latency and access guarantees. A media-delivery token is a short-lived credential for requesting a permitted image variant. Allowing delivery under that credential until expiry reduces repeated permission-service calls, but creates the revocation window stated below.
- Feed latency: Feed metadata p95 below 200 ms inside the serving region.
- Availability: 99.9% eligible feed success inside the serving region.
- Media latency: Preview time to first byte below 300 ms from an available nearby delivery cache; measure it separately from feed metadata.
- Processing time: 95% of valid ordinary images become READY within 30 seconds under planned peak load.
- Feed freshness: Ordinary-author propagation should usually finish within five seconds. Expose backlog; a slightly older authorized feed is acceptable.
- Durability: Acknowledged original storage and metadata commits survive one node or availability-zone failure through correctly placed/configured replicas.
- Regional recovery: Initially use asynchronous replication with an explicit measured recovery point and a one-hour restore target. Derivatives and feeds are rebuildable; lost originals cannot be recovered from feed IDs.
- Private-media revocation: Delivery tokens last at most 60 seconds. An already issued token may remain usable for that interval; immediate revocation requires current authorization checks on every edge request.
Publication and authorization rules
| Rule | Required behavior |
|---|---|
| READY publication | Reference verified durable variants and a retained original. |
| Metadata access | Check current deletion state and membership. |
| Authority partition | Reject new private grants; a cached feed is not permission. |
| Immediate media revocation option | Pay the added edge latency and permission-service dependency for per-request checks. |
The failure contract is separate from the availability percentage. Proper replication supports the stated node/zone guarantee; it does not justify a claim of “100% reliability.”
04Capacity estimates
Workload assumptions and arithmetic
Assume 500 million registered users, one million daily active users, two million photos/day and 200 KB average originals. Add ten feed opens per active user/day with twenty 50 KB previews per page. These are explicit exercise assumptions, not observed traffic. 2M / 86,400 = 23.1 uploads/s; fivefold peak is about 116/s. Feed requests average 10M / 86,400 = 116/s, peaking near 579/s.
Worked estimates
| Quantity | Calculation | What it changes |
|---|---|---|
| Original bytes | 2M × 200 KB = 400 GB/day |
Bulk storage must grow independently |
| Ten-year originals | 400 GB × 365 × 10 = 1.46 PB |
Retention dominates long-term bytes |
| Preview delivery | 10M × 20 × 50 KB = 10 TB/day |
Geographic caches reduce origin traffic |
| Mean preview egress | 10 TB / 86,400 ≈ 116 MB/s |
Delivery is much larger than API payload |
| Follow edges | 500M × 500 × 16 B = 4 TB raw |
Both directions and indexes add cost |
| Metadata illustration | 2M × 284 B × 365 × 10 ≈ 2.07 TB |
Count bytes consistently before indexes |
Capacity implications and limits
The 284-byte record-size estimate is illustrative; actual IDs, strings and indexes change it. Derivatives, replicated copies and backups are additional. If an ordinary author has 300 active followers, 23.1 average photo publications/s yield about 6,930 inbox inserts/s before celebrity exceptions. At peak, about 34,800/s. A single fifty-million-follower author breaks this average immediately. A content delivery network (CDN) that serves 90% of requested image bytes from its caches reduces the 10 TB/day preview origin demand toward 1 TB/day, but viewers still receive 10 TB/day and caches still incur that delivery cost.
For a minimal user-record estimate, assume 500 million users × 68 bytes = 34 GB of raw fixed fields. That is an arithmetic floor, not the size of a production user table: variable profile fields, indexes, access-control records and replicas add bytes. It reinforces why media capacity and metadata capacity need separate estimates.
05APIs and contracts
Request and response example
The author calls POST /v1/photo-uploads with request key upload-90 and {"title":"Sunrise","bytes":200000,"visibility":"followers","checksum":"H900"}. The response names photoId:p900, uploadId:up900, a generation-specific object target and an expiry. The upload capability is issued only after authenticating the owner and is limited to this object generation and operation. A presigned URL is a bearer credential, not a later proof of the uploader’s identity. Bind supported length/checksum conditions cryptographically or through the upload policy, and recheck the accepted object at completion. Completion is a separate authenticated request; possessing a storage upload token cannot publish metadata directly.
Interface contracts
| API | Result and error behavior |
|---|---|
POST /photo-uploads/up900/complete |
Verify object; 202 with processing state; duplicate completion reuses state |
GET /photos/p900/status |
Owner sees uploading/processing/ready/failed |
PUT /following/u17 |
Idempotently follow the author, or create a pending request for private accounts |
GET /feed?cursor=<token>&limit=20 |
Visible metadata, short-lived media tokens, next cursor |
GET /users/u17/photos?before=<time,id> |
Author/time gallery page |
GET /photos/search?q=sunrise&cursor=... |
Title-index candidates filtered by current visibility |
DELETE /photos/p900 |
Owner-authorized tombstone; repeated delete is harmless |
Validation and response semantics
The feed cursor refers to a bounded candidate snapshot and offset or stable ranking key; it is scoped to the viewer and expires. Deletions may make a page shorter, so the server may fetch extra candidates within a work limit. Invalid formats return 400, oversized files 413, request-key payload mismatch 409 and temporary capacity failures 429/503. A client never invents a new upload identity merely because a completion response timed out.
06Data model and access patterns
The model separates the authoritative photo from the work needed to publish and distribute it. A manifest is the list of accepted image variants and their storage references; readers use it to select complete outputs. The outbox stores processing or publication work in the same database transaction as the photo change, while Feed stores rebuildable per-viewer references rather than image bytes.
| Record and fields | Responsibility / constraint |
|---|---|
Photo(photoId, ownerId, createdAt, title, visibility, state, sourceGeneration, originalKey, originalVersionId, manifest, version) |
The manifest identifies preview and large-image objects for one published generation. |
Upload(uploadId, ownerId, requestKey, photoId, checksum, leaseUntil) |
Owns upload retry state. |
Follow(followerId,authorId,status,version) |
Owns relationships. |
Outbox(eventId,photoId,generation,type) |
Records durable work. |
Feed(viewerId,photoId,sortKey) |
Derived; unique by viewer/photo identity. |
The author's gallery query needs (ownerId,createdAt DESC,photoId DESC), not merely an index on photo ID. The viewer's follow list needs (followerId,authorId); publication fanout needs the reverse (authorId,followerId) access path. The title index is another derived view; it cannot authorize a private photo. A gallery query reads WHERE ownerId='u17' AND (createdAt,photoId)<(:t,:id) ORDER BY createdAt DESC,photoId DESC LIMIT 20.
Partition primary photos by a hash of photo ID when independent growth requires it. Maintain an author/time index to avoid querying every photo shard for the author's gallery. Partition inboxes by viewer, with bounded recent retention; shard huge follower lists into pages. Time buckets can reduce old-data scanning but should be combined with hashing or author keys so the newest bucket does not become the only write target. Metadata and outbox changes for a photo share one authoritative transaction. Cross-partition feed inserts are asynchronous, never part of the publication commit.
07Basic working design
Synchronous publication on one server
Start with one app and a SQL database containing photo metadata, follows and original/preview bytes for a small corpus. The author uploads, the app validates and generates the two required variants synchronously, then commits the photo and bytes together. Only after commit does p900 become visible. The upload response can be slow, but the single commit boundary is easy to understand. A failed decode produces no ready photo.
Pull-on-read feed assembly
When the viewer opens a home feed, read the viewer’s followed-author list, query recent photos from those authors, filter visibility, sort by time and return the first twenty. Fetching a bounded number of recent rows per author is a straightforward implementation. The author's gallery is one indexed query. Title search can initially use a modest database text index rather than a dedicated search cluster.
What the baseline buys and where it stops
For a tiny active population this is operationally attractive: one backup covers metadata and original bytes, one transaction publishes, and debugging p900 is local. The baseline's weaknesses are synchronous image processing, centralized media egress and repeated multi-author feed work. It is intentionally a working product rather than an unfinished drawing. Later changes must still return complete, authorized photos while reducing image transfer and repeated feed assembly in requests.
One database commits the image and metadata; feed work repeats across followed authors.
Read each connection in order
- sync1. Upload or open feedUploader and viewer clients → Photo and feed application
- sync2. Commit validated photoPhoto and feed application → SQL photos, follows and image bytes
- sync3. Query followed authorsPhoto and feed application → SQL photos, follows and image bytes
- sync4. Return page and image bytesPhoto and feed application → Uploader and viewer clients
08Find the baseline flaws
| Bottleneck / counterexample | Evidence and design consequence |
|---|---|
| Repeated feed candidate work | Suppose the viewer follows 500 authors and the baseline retrieves each author's latest 100 photos before ranking. That is 50,000 candidates to produce twenty results. At 579 peak feed requests/s, repeating this policy examines roughly 29 million candidate rows/s before loading metadata for those photo IDs. Even a well-indexed table cannot erase that repeated work. The first improvement is to bound or precompute candidates, not merely add a ranking service. |
| Upload/read resource contention | Uploads create another bottleneck. At 116 uploads/s and four seconds of transfer/processing time, roughly 464 uploads are active. Sharing a hypothetical 500-connection or worker budget with feed reads can starve the read path. The exact limit is implementation-dependent; the point is to measure occupancy and isolate work with different duration, rather than assert every server has the same universal connection limit. |
| Premature READY publication | A correctness counterexample appears after naive asynchronous resizing. The API inserts READY metadata and pushes a job, then the worker crashes before writing the preview. The viewer sees a broken feed item. Or the database commits PROCESSING but the separate queue send fails, leaving the photo stuck forever. These are different failures. To prevent the broken photo, a worker verifies all required images before committing their manifest. To prevent lost work, store the processing event with the state change in a transactional outbox and send it after commit. |
09Improve the design, step by step
Move original bytes to private object storage and separate upload/read pools. The trigger is growing byte retention plus hundreds of slow uploads. Direct constrained uploads remove bulk transfer from feed servers; dedicated completion APIs verify objects. This improves read isolation and independent storage scaling. It costs two-store coordination, token management and orphan cleanup. Keeping bytes in the database remains simpler for small corpora; do not split before the operational benefit is real.
Process variants asynchronously with a durable outbox. The trigger is decode latency and variable image complexity. Commit PROCESSING with an event, then workers create immutable generation-specific variants and atomically publish a complete manifest. This gives predictable upload acceptance and recoverable work. It costs queues, worker capacity, duplicate handling and a visible processing state. Synchronous processing is preferable when bounded small inputs reliably fit the response budget.
Introduce hybrid feed preparation. The trigger is repeated 50,000-candidate assembly. Ordinary authors' photo IDs are inserted into active followers' inboxes; celebrity photos stay in author lists and merge on read. This makes ordinary feed reads cheap while avoiding fifty million writes for one celebrity upload. It costs inbox storage, fanout checkpoints and two-path deduplication. Pure pull is attractive for inactive viewers or small follow lists; pure push works when follower counts and write amplification remain bounded.
Add media delivery caches, metadata partitions and replicated authority. The triggers are roughly 10 TB/day preview delivery, retained metadata growth and zone-failure durability. Immutable variants are cached near viewers; logical partitions distribute records and queries; replicated leaders protect acknowledged changes. Costs include origin-fill bursts, routing epochs, duplicate storage and authorization-token lifetime. A single larger replicated database remains viable until measurements justify partitions; SQL is not ruled out by the product's name.
Each change preserves the invariant that only a committed READY manifest enters candidate feeds. None permits a fanout worker or CDN to decide independently that private content is public.
10Detailed architecture
Upload control and metadata authority
The edge routes upload-control traffic to an upload API and feed/search traffic to a read API. The phone sends bytes directly to private object storage using a constrained upload target. The metadata database stores uploads, photos, follows and publication events. Each partition’s replicated leader commits changes in order. The diagram shows representative groups, not one global lock for every photo.
Processing and feed workers
An outbox relay forwards processing and publication events to a durable work queue. Image workers write immutable variants, then ask the database partition storing the photo to publish its manifest. Feed workers consume publication events and page through follower lists into viewer inboxes. Search indexing and author-list updates consume the same committed publication stream, with idempotent event identity. They may lag without changing p900's authoritative state.
Authorized feed and media reads
Read APIs combine inbox candidates with celebrity author lists, batch-load metadata, check current visibility/membership, and return a bounded page. A separate media edge validates short-lived access tokens and serves CDN-cached bytes or fetches the private origin. Thus authorization is in front of delivery, not an optional caption beside a public bucket. Ranking may degrade to authorized recency order; permission checks may not degrade to “allow all.” Logical partition routing is versioned, and migrations fence old owners before accepting writes at new locations.
Only a committed READY manifest emits the publication event. Inbox and search entries are candidates; the read path still authorizes the photo.
Read each connection in order
- sync1. Upload control / page requestMobile and web clients → Edge request routing
- sync2a. Admit upload sessionEdge request routing → Upload control API
- sync2b. Request visible candidatesEdge request routing → Feed, gallery and search API
- sync3. Reserve / pin verified source versionUpload control API → Partitioned metadata authority
- sync4. Upload scoped original g1Mobile and web clients → Private original and variant store
- replication5. Replicate authoritative statePartitioned metadata authority → Metadata replicas
- async6. Relay committed outboxPartitioned metadata authority → Outbox relay and work queue
- async7. Process current attemptOutbox relay and work queue → Image processing workers
- sync8. Write immutable variantsImage processing workers → Private original and variant store
- sync9. Guarded READY + outboxImage processing workers → Partitioned metadata authority
- async10. Publish references and indexesOutbox relay and work queue → Feed and index workers
- async11. Idempotent candidate updatesFeed and index workers → Inbox, author and title indexes
- sync12. Merge bounded candidatesFeed, gallery and search API → Inbox, author and title indexes
- sync13. Load photo metadata and check accessFeed, gallery and search API → Partitioned metadata authority
- sync14. Page and scoped media tokenFeed, gallery and search API → Mobile and web clients
- sync15. Present token + bound sessionMobile and web clients → Authorized media edge and CDN
- sync16. Cache miss: private originAuthorized media edge and CDN → Private original and variant store
11Write path and acknowledgement
The upload protocol publishes only a verified image manifest and records downstream work durably. Photo p900, upload up900 and request upload-90 provide concrete identifiers for the transitions.
- The author authenticates as u17 and creates upload up900 using upload-90. A transaction reserves photo p900 in UPLOADING with a checksum, generation and deadline.
- The upload client uploads
originals/p900/g1. The storage token cannot write other users' keys. A failed transfer resumes/retries the same session rather than creating an unrelated photo. - Completion verifies expected size/checksum and permitted image format, then commits PROCESSING and outbox event
process-p900-g1. Return 202: original accepted, publication pending. - A relay publishes the event. Worker W1 decodes with pixel/dimension limits, removes disallowed metadata such as location data when policy requires it, and writes preview and large variants under an attempt-specific immutable prefix.
- W1 verifies every required output and submits a manifest plus its claimed job generation. The metadata owner atomically changes PROCESSING to READY only for the current valid generation and inserts
ready-p900-g1in the outbox. - Feed workers read this committed event, page through ordinary active followers, and insert
(viewer31,p900)if absent. Title search and author/time indexes receive idempotent updates. - The author's status request returns READY. If completion or publication responses were lost, retry reads the existing session/photo state. Nothing in the protocol requires generating a second p900.
Unused attempt outputs never appear in the manifest. Delete them only after the metadata state prevents any current worker from publishing them.
A generation-shaped key is not automatically immutable in an object store. S3 presigned upload URLs can be reused before expiration and can replace the current object at the key. One safe implementation uses a versioned bucket, verifies an exact accepted VersionId and checksum at completion, stores that VersionId in the photo record, and makes every worker read that exact version. Later uploads to the same key cannot change the source already accepted. Alternatively enforce a conditional create-only upload with the required signed checksum. Retain the accepted version through lifecycle rules; default current-key GETs and blanket noncurrent-version expiration would break the version-pinned design.
12Read and delivery path
Feed assembly chooses candidate IDs, checks current access, then issues bounded media grants. A twenty-item request illustrates these responsibilities without treating an old inbox entry as permission.
- The viewer authenticates and requests a twenty-item feed page. The read API validates the cursor's viewer identity and snapshot lifetime.
- It loads a bounded inbox window and recent photos from followed high-fanout authors. It merges by sort key, deduplicates photo IDs and applies a candidate limit to keep one request's work bounded.
- Batch-fetch photo metadata by ID, using caches for immutable fields while rechecking authoritative deletion/visibility and current membership according to the private-access contract. Remove p900 if the author deleted it or the viewer no longer has access.
- Rank eligible candidates by recency and bounded relevance signals. Return twenty results or a shorter page with a cursor if the bounded candidate window contains too few eligible items. Do not issue unbounded fanout queries just to fill every page perfectly.
- For p900, issue a media token scoped to the viewer or the authorized session, object generation, variant and 60-second expiry. Return the title, owner, dimensions and preview route.
- The viewer's device requests the media edge. It validates the token before serving cached preview bytes; a miss fetches the private origin. The original remains inaccessible unless separately authorized.
Polling, long polling or push notifications can tell the viewer that new items are available. These delivery mechanisms are separate from database fanout-on-write. Coalesce notifications for busy followers rather than pushing one user-interface refresh for every photo. Gallery and title-search reads use their own indexes but finish with the same visibility checks.
13Correctness deep dive
Retries need a publication guard
A worker can finish writing image variants, then lose its queue acknowledgment. The queue may therefore send the job again. “At least once” means a job can be delivered again after a timeout; it does not mean the photo should be published twice. Use a job generation and a lease token whose validity is checked by the metadata authority during publication.
| Operation | Authority check | Durable effect |
|---|---|---|
| Claim processing | Photo PROCESSING, no active valid claim | Save attempt token 41 and lease |
| Reclaim after expiry | Token 41 expired, still PROCESSING | Save token 42; old token becomes invalid |
| Publish manifest | Current token matches, lease valid, photo not deleted | READY plus one unique publication outbox event |
| Repeat publication | Same committed generation already READY | Return existing manifest, no second event |
| Insert feed reference | Unique (viewerId,photoId) absent |
One candidate reference; duplicate is harmless |
Stale-worker publication proof
Fanout checkpoint and deduplication proof
Attempt-specific objects prevent stale byte writes; the metadata token comparison prevents stale publication.
Read each connection in order
- syncClaim token 41Worker W1 → Photo authority
- syncWrite attempt-41 previewWorker W1 → Object store
- syncAfter expiry: claim token 42Worker W2 → Photo authority
- syncWrite complete attempt-42 setWorker W2 → Object store
- syncPublish manifest under token 42Worker W2 → Photo authority
- returnREADY plus publication outboxPhoto authority → Worker W2
- syncLate publish under token 41Worker W1 → Photo authority
- returnReject obsolete tokenPhoto authority → Worker W1
14Failure and recovery
| Failure / trigger | User outcome, surviving state and recovery |
|---|---|
| Image worker crash | Original g1 and PROCESSING state survive. Another worker claims a new attempt after the lease expires, regenerates variants and tries the guarded publication transaction. The author sees processing longer; the viewer sees no broken READY entry. If decoding repeatedly fails, mark a durable failed state and expose a useful error rather than retrying forever. |
| Metadata zone failure or partition | A surviving majority can elect a leader and preserve committed READY manifests. A minority cannot publish or grant new private access. Existing short-lived media tokens remain valid until their declared expiry; after that the edge cannot mint replacements without authorization. A full-region outage has the separately stated recovery window. CDN copies do not replace backups of originals or ownership metadata. |
| Celebrity burst | One fifty-million-follower post must not enqueue fifty million urgent writes onto the ordinary path. Classification sends it to the author-list path; read caches share popular metadata and preview bytes. If ordinary fanout backlog grows, prioritize active viewers and maintain a bounded catch-up window. Return an older authorized feed with a freshness indicator while publication progresses. |
| Ranking or search outage | Feed reads fall back to authorized recency candidates; title search may return a temporary error rather than leak unfiltered cached results. Deletion first tombstones metadata, then asynchronously removes indexes and media. Old inbox/search entries are harmless references only because read-time policy is enforced. Already downloaded images remain beyond the service's revocation control. |
15Operations, security, and cost
Processing, delivery and feed signals
Track upload-to-ready percentiles, oldest processing lease, invalid-image rate, whether every manifest references existing variants with matching checksums, and whether backups restore the original images successfully. Feed metrics include p95/p99 metadata latency, candidate count, per-author fanout work, publication lag and fraction of requests using celebrity merges. Delivery metrics include byte-hit ratio, origin bandwidth, token failures and denied private requests. A simple API success counter would miss most of these user-visible failures.
Storage, egress and inbox cost
Cost depends on how long originals are retained, the number and sizes of derived image variants, bytes delivered to viewers, and the number of follower-inbox references written per photo. If each candidate reference occupies an illustrative 32 bytes, 300 follower references cost 300 × 32 = 9.6 KB per ordinary photo, versus a 200 KB original. Fifty million references cost 1.6 GB for one celebrity photo before indexes and replicas. That arithmetic explains a hybrid policy better than an arbitrary celebrity label. Determine the push/pull threshold from expected active follower reads during the useful feed window and measured merge cost.
Untrusted image handling
Exercise malformed image headers, decompression bombs, oversized dimensions, expired upload tokens and forbidden object paths. Remove unnecessary location metadata according to product policy. Do not log private media tokens. Enforce object-store access policies and account ownership at completion, not only when upload starts.
Migration and failure drills
For a partition migration, copy records and author indexes, replay changes, compare sample queries, fence the old epoch and cut over routing. Test W1/W2 lease races, a repeated fanout page, deleted photos in cached feeds, and restoration of originals with manifests. Roll out ranking separately from correctness-sensitive visibility filtering so a model change cannot bypass access checks.
16Decision ledger and limitations
Partitioning chooses which storage group owns a photo or an index entry. Feed preparation chooses when references are copied or merged for viewers. These are independent decisions: balancing primary photo records does not by itself make an author's gallery query local or bound a celebrity's fanout work.
| Choice | Benefit | Cost and consequence | Change trigger |
|---|---|---|---|
| Hash photo-ID primary records | Spreads different photos and bytes | Gallery requires author/time index | Owner-local transactions dominate access |
| Owner-based partition | Gallery locality | Prolific/hot authors skew one owner | Split large owners across time buckets |
| Hybrid feed candidates | Cheap ordinary reads, bounded celebrity writes | Two paths, checkpoints and deduplication | Workload shifts toward mostly inactive viewers |
| Immutable variant manifests | Safe retry and cache identity | Orphan attempts and retained originals | Strongly transactional media storage simplifies it |
| Short-lived private media tokens | Delivery edge avoids central check per byte request | Revocation bounded by token lifetime | Immediate revocation becomes mandatory |
Time-sortable photo IDs can include timestamp, generator identity and per-tick sequence. Each generator must handle clock rollback and sequence exhaustion without issuing duplicates. An ID format with a 31-bit seconds field has a finite time horizon, and its 9-bit sequence permits only 512 IDs per second within one allocation scope, so average 23/s does not justify safety under peaks or multiple generators. Use a proven larger scheme or allocated IDs with explicit authority.
Disjoint odd/even database sequences are another allocation alternative. Their ranges must remain disjoint through failover; standby promotion cannot reset a sequence and reuse values. A logical partition map is more flexible than hard-coded id % currentServerCount, but moving it requires a fenced migration, not just editing a configuration file. LRU metadata caching is reasonable when measured locality supports it; popularity and byte size may justify admission limits beyond recency alone.
17Interview closing
“I designed photos, follows, galleries, title search and a feed with durable publication and authorized delivery. The assumptions give about twenty-three average uploads per second but 400 GB of new originals and ten terabytes of preview delivery per day. I start with a working transactional version, then move bytes and resizing out of feed servers, add a durable outbox, and prepare ordinary followers' candidate lists.
“The hard guarantee is that a READY photo references a complete verified manifest, and a stale worker cannot replace the current generation. Feed delivery is eventually updated and idempotent; it never grants access by itself. I use hybrid fanout because a fifty-million-follower author makes per-follower writes unreasonable. Media caches serve immutable variants, while authorization happens before delivery and private tokens have an explicit sixty-second lifetime.
“I accept bounded feed staleness and some extra index complexity. The next measurements are candidate-merge cost, fanout backlog for active users, and origin bytes after cache loss.”
If the interviewer changes the feed to a highly personalized ranking, keep candidate generation and current visibility as separate stages. Add a versioned ranking model over bounded eligible candidates, evaluate relevance and latency, and stabilize pagination. Do not let a ranking score become evidence of permission or replace the publication invariant.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
What are the three different things you store for one photograph?
Reveal a model answer
“The original bytes, metadata describing ownership and state, and references in author or follower lists. Derivatives are rebuildable versions of the original; feed entries are candidate references. Losing each has a different recovery story.”
Interviewer follow-up
Which one is the CDN responsible for?
Reveal the follow-up answer
“Delivering image variants. It is not the authority for ownership, publication state, or who the viewer follows.”
What the answer must demonstrate: Do not store full images in every follower feed.
A viewer follows 500 authors. How would you assemble a twenty-item photo feed, and when would you precompute it?
Reveal a model answer
“Initially I query recent author lists and merge a bounded set. If repeated reads make that too expensive, ordinary authors distribute IDs into active followers’ inboxes. I merge celebrity-author lists on read and filter current permissions before ranking.”
Interviewer follow-up
Why not push every celebrity photo?
Reveal the follow-up answer
“Fifty million followers means fifty million feed writes for one post, including many inactive readers. Pulling that author’s list during actual reads avoids much unused work.”
What the answer must demonstrate: Explain fanout’s unit of work before choosing it.
The resize worker crashes after producing one image variant. What happens?
Reveal a model answer
“The photo remains PROCESSING, with its original retained durably. The replacement reads the exact verified source version recorded at completion but writes to its own attempt-specific output paths. It verifies every required variant, then atomically checks its current worker token before publishing READY plus the feed outbox event. The old attempt cannot overwrite the accepted paths or win publication after its token is replaced.”
Interviewer follow-up
Why not expose the original immediately?
Reveal the follow-up answer
“That is a possible product choice, but I would specify the fallback and its size/security implications. This contract waits for a valid preview to avoid broken or huge feed images.”
What the answer must demonstrate: An output object alone must not imply ready metadata.
Why does a photo service need an author/time index in addition to a PhotoID primary key?
Reveal a model answer
“A PhotoID lookup retrieves one known photo. A profile asks for an author’s newest photos, so it needs an owner/time access path, such as (ownerId, createdAt, photoId). Hashing primary records otherwise scatters that range query. I add the index because of the query shape, not because photo IDs are insufficiently unique.”
Interviewer follow-up
Does a timestamp-bearing ID solve the owner query?
Reveal the follow-up answer
“It helps order known IDs, but does not group one owner’s IDs. I still keep an author list or suitable secondary index.”
What the answer must demonstrate: Ordering and locating are different tasks.
A private photo is deleted after its ID entered a viewer’s feed cache. How do you prevent the stale candidate from disclosing it?
Reveal a model answer
“An inbox entry is only a candidate. Before returning its metadata or issuing a media grant, I check current deletion and membership at the authority. Asynchronous cleanup removes stale references but is not the permission boundary. Previously issued private media tokens remain usable for up to the stated 60 seconds; immediate revocation would require current checks at the delivery edge.” The edge must also check any claimed viewer/session binding against the authenticated requester; verifying a token signature alone does not enforce that binding.
Interviewer follow-up
Can eventual feed freshness also apply to permission changes?
Reveal the follow-up answer
“Not automatically. A delayed new photo may be acceptable while delayed revocation is not. They need different freshness guarantees.”
What the answer must demonstrate: Do not use one consistency slogan for every read.
What does a 10 TB/day preview estimate tell you?
Reveal a model answer
“Media delivery is a separate bandwidth path. I consider derivative size and distributed caching, then measure origin byte-hit ratio. It does not mean metadata needs the same capacity or that a CDN eliminates viewer traffic.”
Interviewer follow-up
What storage estimate is still missing?
What the answer must demonstrate: Separate payload estimates from replicated capacity.
A paused resize worker wakes after a replacement published. What stops it corrupting the photo?
Reveal a model answer
The metadata owner checks the worker token atomically when changing PROCESSING to READY. The replacement has a new token, so the old worker cannot publish. Crucially, each attempt writes immutable object names; otherwise the stale worker could overwrite accepted bytes even if its metadata update were rejected.
Interviewer follow-up
Can you delete the old attempt immediately?
Reveal the follow-up answer
Only after the authority proves it cannot still publish and the producer is bounded or fenced. Then its objects are unreferenced candidates for idempotent reclamation. A guessed timeout without a publication guard is unsafe.
What the answer must demonstrate: Reject the old worker’s manifest update and prevent it from overwriting the accepted image objects.
Show the cost that makes a hybrid feed worthwhile.
Reveal a model answer
At 300 active followers, one 32-byte reference per follower is about 9.6 KB per photo. At fifty million followers it becomes 1.6 GB before indexes and replicas. I keep that large author’s recent list and merge it on active readers’ requests, while ordinary authors benefit from prepared inboxes.
Interviewer follow-up
Does that mean clients must use polling for celebrities?
Reveal the follow-up answer
No. Database candidate preparation and client notification are separate decisions. I can push a coalesced “new items available” signal while still merging the celebrity list on read.
What the answer must demonstrate: Do not confuse fanout-on-write with WebSocket or push notification transport.
Blank-page exercise · 45 minutes
Build the answer yourself
Design photo uploads, galleries and a follower feed. Derive the storage and delivery workload, evolve from a single-server baseline, then handle a 50-million-follower author and a resize worker that resumes after its replacement publishes.
- Name original, derivative, metadata, and feed reference.
- Calculate uploads, retained originals, and preview delivery bytes.
- Show p900 state transitions and the durable event handoff.
- Compare fanout-on-write/read with an actual follower count.
- Keep author/time retrieval and photo lookup distinct.
- Explain deleted-photo behavior despite a stale inbox.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design a photo-sharing serviceWhat does fanout-on-write copy?Recall first, then reveal
Photo references into eligible followers’ candidate lists, not complete image bytes.
Many inboxes; one original.
Return to lessonDesign a photo-sharing serviceWhy keep ownerId + time after hashing photo IDs?Recall first, then reveal
Because fetching one author’s recent photographs is a different access path from finding one photo by ID.
Identity finds one; index finds a list.
Return to lessonDesign a photo-sharing serviceA preview object exists while photo metadata is PROCESSING. May the feed expose it?Recall first, then reveal
No. The worker verifies all required outputs and commits the guarded READY manifest before emitting the publication event.
A file is not a published photo.
Return to lessonFinal revision
Summary and interview notes
A photo service publishes verified media manifests and distributes candidate references, then authorizes metadata and byte delivery separately. Hybrid feeds reduce ordinary read work without turning one popular author into tens of millions of mandatory inbox writes.
Remember these points
- At the stated workload, originals add 400 GB/day while previews deliver 10 TB/day; storage and delivery need separate capacity plans.
- Pin the exact verified source version so a reusable upload credential cannot change an accepted image.
- A current worker token and attempt-specific immutable output names protect both manifest publication and external bytes.
- Copy ordinary authors’ photo IDs into active followers’ lists. Merge celebrity lists during reads, limit candidates, and save progress only after inserts are safe to repeat.
- A stale inbox is not permission, and viewer-bound media tokens require the edge to verify the matching identity.
Interview tips
- Show why 500 authors × 100 photos is expensive before introducing prepared inboxes.
- Resume an old resize worker after a replacement publishes; protect both the pointer and object names.
- Ask whether an upload URL can be replayed and whether a media token is bearer-only or actually identity-bound.
Important qualifications
- S3 presigned URLs are reusable bearer capabilities until their applicable expiry; a named generation alone is not immutable storage.
- The 60-second media-token contract allows a bounded revocation delay; immediate revocation needs a different serving check.
- A derived gallery/title index may lag; it must still filter current publication and privacy state.
Technical references
- S3 presigned upload URLsExplains constrained direct-to-storage uploads; image publication and authorization remain application responsibilities.
- PostgreSQL multicolumn indexesDocuments why an index’s column order matters for the owner/time access path.
- Transactional outbox patternSupports a reliable publication-to-processing/feed handoff.
- Amazon S3: Retrieving Object VersionsExact version retrieval pins the accepted original even if a presigned URL later writes a newer version.
- Amazon S3: Conditional WritesAlternative create-only enforcement; an object-key naming convention alone does not prevent overwrite.
Practice marks stay in this browser.