Concept lesson · Foundations
System design interview framework
Start here
Definition
System design is the process of defining a service’s components, data, interfaces and interactions so it meets stated functional and quality requirements. A system-design interview asks you to explain and defend those choices under an explicit workload and failure model.
Why it matters: A feature description says what a user wants; a design explains which component performs each step, where the facts are stored, and what happens if a step fails.
State the requirements, trace a working request, then justify each design change using a capacity limit or failure it must handle. Explain both the benefit and the cost.
Read the diagram step by step
- For a file-sharing service, define upload and download requirements. The publication contract permits downloads only after a complete file is ready.
- Estimate download bytes and metadata separately, then trace upload, completion and abc123 lookup.
- Large file bytes justify object storage; the new publication gap requires uploading and ready states.
- Prove retry after a lost completion response, then close with the invariant, bottleneck and next test.
Worked example
Upload record 42 refers to a 2 MB PDF. The application stores state uploading, verifies the complete file, then changes the record to ready; downloads require a ready, permitted file.
Key takeaways
- Requirements determine the design.
- Trace one complete request and its durable result.
- For each change, explain the benefit, cost and failure behavior.
You will learn to
- Turn an ambiguous prompt into agreed user actions, measurable targets, and rules the design must preserve.
- Draw and explain one complete request before scaling.
- Answer follow-ups by changing the design and stating the cost.
- Distinguish scaling an application from splitting it into independently deployed services.
Practice in this chapter
9 interview questions with model answers and follow-ups.
Go to interview practiceWorkload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01System design: definition and purpose
System design is the process of defining components, data, interfaces and interactions that satisfy a service’s requirements. A component is a part with a specific job: an application accepts a request, a database stores searchable records, and file storage holds uploaded bytes. An interface is the agreement for asking a component to do work. The interview asks you to show how these parts cooperate and how the design behaves when traffic increases or a part fails.
You are not expected to guess a company's private architecture. You are expected to build a plausible design under stated requirements. A requirement states what the product must do or how well it must work. A tradeoff is a benefit gained by accepting a cost or limitation elsewhere. For example, keeping a second copy of an uploaded photo helps survive a disk failure, but uses more storage and requires a rule for when the upload is considered safe.
A complete answer covers requirements, workload estimates, API contracts, data ownership, component responsibilities, and failure behavior. Use one operation to verify that the proposed components form a working system. A bounded example throughout the method is a file-sharing API for PDF worksheets, with upload, download, and deletion; record 42 identifies a 2 MB PDF.
02Functional and non-functional requirements
Start by identifying actors, operations, and access rules. Clarify account requirements, file-size limits, link expiration, and revocation. These decisions determine the API and how current its permission checks must be: a public permanent link needs different read checks from a link that must stop authorizing downloads immediately after revocation.
Functional requirements: what the product does
- Upload. An authenticated uploader can create an upload for a PDF up to 10 MB.
- Share. Create an unlisted download link, with optional expiration. “Unlisted” means the link is hard to guess but possession permits access; private group sharing needs an additional authorization policy.
- Download. Retrieve a permitted file through its link.
- Delete. The owner can revoke future downloads.
These capabilities need file storage and link metadata. Collaborative editing and document grading are outside this example’s scope.
Terms used in the targets
Here, latency is the time one request takes; p95 is a threshold met by approximately 95% of the measured requests. Availability measures whether the requested operation can be used successfully. Authoritative means the component whose recorded decision is treated as truth. These terms make a requirement testable instead of merely saying “fast and reliable.”
Non-functional requirements: how well it must work
- Performance. Start an allowed download within 200 ms at p95. Measure lookup and first-byte delay separately.
- Correctness. A deleted link cannot start a new download. Check the authoritative deletion policy.
- Availability. Target 99.9% successful eligible download attempts per month. Define the measurement and a failure plan.
The numeric targets here are assumptions for practice. State them, then invite the interviewer to change them.
03Workload and capacity estimates
Assume 100,000 uploading accounts, two files per account per month, a 2 MB average file, and 30 downloads per file. Use the numbers to identify the dominant work:
| Step | Calculation | Design implication |
|---|---|---|
| 1. New file data | 100,000 × 2 × 2 MB = 400 GB/month |
Retention determines accumulated storage |
| 2. Delivered bytes | 400 GB × 30 = 12 TB/month |
Downloading bytes dominates uploading bytes |
| 3. Average downloads | 6,000,000 / 2,592,000 seconds ≈ 2.3 requests/s |
An average hides concentrated bursts |
| 4. Peak downloads | Assume a measured or interviewer-supplied 1,000 downloads/s | Use the peak before deciding server capacity |
These figures justify separating file delivery from metadata requests. They do not establish that the metadata database needs hundreds of shards.
Interview checklist:
- Show units. “400 GB” is storage; “400 GB/month” is growth; “1,000 requests/s” is a rate.
- Ask about retention. Do this before multiplying monthly growth into lifetime storage.
- Estimate to decide. Do not spend ten minutes estimating a number that will not change the architecture.
04APIs, data model and source of truth
An API is the agreement between a caller and a service. For example, use POST /worksheets with a filename, size, and expiry to create an upload session. A response returns worksheetId=42 and an upload destination. A completion call validates that the file exists before changing its state to ready. GET /links/abc123 looks up a download, and an authorized DELETE /worksheets/42 revokes future access.
Follow an upload through the API, database and object store. Mark ready only after upload validation.
Remember: The database locates the file; the object store holds its bytes.
Read the diagram
- Trace the separate metadata and byte paths from one client.
- The API saves upload U7 and its object key in the metadata database.
- The client uploads bytes to object storage; completion validation permits state ready.
Try from memoryWould copying the metadata row copy the uploaded file?
No. The row contains an object reference and state; copying the file requires copying its bytes.
Metadata is information about a file rather than the file’s own bytes. Keep one metadata row: Worksheet(id, ownerId, objectKey, objectVersion, state, expiresAt, deletedAt). The bytes live separately under an object key such as worksheets/42/v1. An object key is the storage address of the file, not the public permission to download it. The query you must support is a point lookup of one link or worksheet, so a primary-key index is a useful first choice.
A primary key uniquely identifies a database row. An index is a maintained lookup structure that helps the database find matching rows without inspecting every row. Here the link token identifies one mapping, and the worksheet ID identifies one metadata record; neither lookup needs to search the file bytes.
The initial implementation can be one application and one database plus durable file storage. Splitting APIs into many services before describing this contract adds complexity without establishing a correct request path.
The public token also needs a stored mapping: Link(token PRIMARY KEY, worksheetId). Generate an unpredictable token for an unlisted link; the short abc123 above is only a readable example, not an adequate security design. A unique owner/request-key record can make upload-session creation retryable. The file row and its request result commit together; repeating that request returns the same session rather than allocating another file.
05Worked example: upload, download and retry
Publishing a worksheet requires agreement between two stores: the metadata database and the file store. The database must not advertise a file whose upload is incomplete. An immutable object version is a particular stored version whose bytes do not change; completion verifies that version and makes the metadata point to it. The following sequence uses states to coordinate that publication.
- An authenticated
POST /worksheetsvalidates the caller’s quota and stores row 42 with stateuploading. - The upload writes bytes to the designated object key. The storage layer must record the complete object; an interrupted transfer does not make it downloadable.
- The completion operation verifies a specific immutable object version, records its checksum and size, and conditionally changes row 42 from
uploadingtoreadyonly if it has not been deleted. A checksum summarizes bytes so corruption can be detected. Atomic means the state transition is indivisible; another request cannot observe half of the row change. GET /links/abc123checks row 42’s readiness, expiry, and deletion state before allowing transfer.- An authorized delete records revocation first. Later cleanup removes unused bytes; delayed cleanup does not reauthorize the link.
The server can save a change successfully even if its response never reaches the client. Handling retries after this failure requires idempotency: repeating one logical operation must recover the same business result. If completion commits but its response is lost, retrying with the stable upload ID returns row 42’s existing ready state instead of creating another file. This failure test identifies where the deduplication record and result must be durable.
Test deletion during upload completion as well as a lost response. If deletion commits first, completion must fail its state check and leave the row deleted. If completion commits first, deletion revokes the published file. A late upload must not overwrite the verified object version. For downloads, the access check determines the order: a transfer authorized before deletion may finish; a check after deletion must reject it. If current metadata is unavailable, block new downloads. HTTP methods and object storage do not enforce these rules on their own.
- 1 → 2create upload sessionUpload request: record 42 → Application: uploading row
- 1 → 3send file bytesUpload request: record 42 → Object store: complete PDF
- 2 → 4verify bytes, publish readyApplication: uploading row → Database: ready row
- 5 → 2check link and permissionDownload request: abc123 → Application: uploading row
- 2 → 3authorize permitted downloadApplication: uploading row → Object store: complete PDF
06Scaling, replication and recovery
When one application cannot handle the measured peak, add stateless application instances and a load balancer, which distributes incoming requests among them. Stateless means another instance can handle the next request because essential worksheet state lives in shared durable storage. If popular public worksheets account for most bytes, a content delivery network can serve permitted cached files close to clients. Revocable/private files need a compatible authorization and cache-expiration design.
Replicate important records, decide what a successful upload promises about durability, and test restoration from backups. Replication means maintaining live copies; a backup lets you recover an older version after a bad deletion. They solve different failures.
In a 45-minute practice session, spend approximately five minutes clarifying, seven on quantities and contracts, ten drawing and tracing, fifteen on the most important bottleneck and failure, and eight reviewing. The interviewer may redirect you. Follow that signal rather than treating the time allocation as a script.
07Monoliths, modular monoliths and service boundaries
A microservices architecture separates capabilities into independently deployable services that communicate through APIs or messages. Each service controls changes to its own data; other services use its contract instead of changing its tables directly. This can help teams release independently and give a demanding component its own resources. It also adds network calls, compatibility work and more components to operate. See Martin Fowler’s discussion of microservice trade-offs.
For a concrete example, the checkout design initially keeps order creation and stock reservation in one database transaction. The application can have separate order and inventory modules without splitting that transaction across services. If browsing grows much faster than purchasing, a separate catalog search service can scale its derived product index while checkout keeps authoritative prices and stock checks together.
Splitting inventory into an independent service needs a stronger reason, such as a shared reservation capability serving several products with its own release schedule. The order and reservation would then commit separately. Define reservation expiry, retries and recovery before claiming the purchase succeeds; use the distributed-workflow lesson for that coordination. Moving code into separate processes does not make the two commits atomic.
The top design keeps order and inventory modules in one deployment with a shared database transaction. The bottom design gives each service its own data; an API call does not commit both databases atomically.
Remember: More application instances do not require more service boundaries.
Read the diagram
- A modular monolith contains order and inventory modules in one deployment; multiple copies of that deployment can run behind a load balancer.
- Its order and inventory updates can share a transaction in the orders and stock database.
- Independent order and inventory services communicate by API or message and each owns its database.
- Separate database commits require a coordinated transaction or a recoverable workflow; the network arrow alone supplies neither.
Try from memoryDoes putting an order module and an inventory module on separate servers preserve their original local transaction?
No. Separate service-owned databases change the transaction boundary. Define a distributed transaction or durable reservation workflow with retries and recovery.
| Decision | Benefit | Cost to explain |
|---|---|---|
| Keep related modules together | Local calls and simpler transaction boundaries | Components share a release and resource allocation |
| Extract a service with a clear responsibility | Independent releases, capacity and ownership | Remote failures, compatible contracts and cross-service recovery |
Choose boundaries around responsibilities that can evolve independently. A diagram box may be a module, a process or a replicated service; say which you mean. Explain how a caller behaves when a service fails, because separation alone does not prevent an outage from spreading. Microsoft’s architecture guidance describes these deployment, data-ownership and failure-handling concerns.
08Interview example: explain a storage choice
Interviewer: “Why not store the PDF in the database?”
Candidate: “The database needs small records for ownership and link lookups. Our estimated 12 TB of monthly downloads is mostly file bytes, so I would put those bytes in object storage and keep their keys in the database. That allows downloads to scale independently. The added problem is publication across two stores; I handle it with uploading and ready states, and make completion retryable.”
This answer contains a choice, a workload-based reason, a new failure risk, and a concrete mechanism. If you cannot explain those four parts for a component, revisit whether it belongs in the first design.
| Choice | Useful when | Added cost or limit |
|---|---|---|
| One application and database | The workload fits and a complete request is easy to explain | One process may limit capacity; recovery still matters |
| Multiple application instances | Application processing or availability is the limit | Essential state must be shared or recoverable |
| Separate object storage | Large files dominate retained or delivered bytes | File publication and metadata need explicit states |
| CDN for repeated public files | Nearby copies materially reduce origin work | Cached authorization must respect the revocation promise |
When answering a follow-up, identify which requirement has changed before adding or replacing components in the diagram. If the interviewer changes worksheets from unlisted to private groups, the concrete new requirement is “only authorized group members may download.” Add a per-reader permission check and explain its failure behavior; the PDF storage itself does not have to change.
09The complete interview sequence
The 36 design chapters expand this method into a full practice interview. Their sections are preparation material: do not recite every paragraph or spend equal time on every stage. In a live interview, establish the complete outline, trace the central request, and use the interviewer's questions to decide which part of the design to explain in greater detail.
| Pass | Sections to reconstruct | Evidence your answer should contain |
|---|---|---|
| Agree on the problem | Scope, functional requirements, non-functional requirements | A specific API operation, acceptance criteria, quantitative assumptions, exclusions and one hard invariant |
| Make a working system | Estimates, APIs, data model, baseline | An actual payload, key/index, query and durable commit point |
| Discover limits | Baseline flaws and ordered improvements | A measured or estimated bottleneck, or an ordering of concurrent operations that breaks a requirement; then a proposed fix, its benefit, its cost and an alternative you rejected |
| Defend the developed system | Detailed architecture, write path, read path, correctness deep dive | How requests reach each component, which component can update each record, when success is acknowledged, where background work begins, and what happens when two clients act concurrently or a component crashes |
| Operate and conclude | Failures, operations/cost, decision ledger, closing | User-visible degradation, surviving state, recovery, remaining limitation, and a concise spoken recap |
Draw the baseline first. When changing it, point to the failed requirement: “This cache removes repeated reads, but introduces up to 25 seconds of stale access, so it violates our original immediate-revocation promise unless we change that contract.” The change is not justified merely because a cache is conventional. A mature answer may keep the simpler design when the requirement does not pay for the added complexity.
Close in roughly 60–90 seconds: restate the requirement, describe the resulting request path, name the invariant and its mechanism, acknowledge the largest cost, and propose the next measurement. Then practise changing one requirement. Changing public file sharing to private group sharing introduces per-reader authorization and revocation; the existing object store remains useful, but the permission decision must change.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
What is system design, and how would you begin “design file sharing”?
Reveal a model answer
“System design defines the components, stored data, interfaces and interactions needed to meet requirements. For file sharing I first ask who uploads, who downloads, size limits, and whether links are public, private, expiring or revocable. Then I agree on volume and what success means, and trace one upload before adding capacity.”
Interviewer follow-up
Should you ask twenty questions before drawing?
Reveal the follow-up answer
No. Resolve the few ambiguities that change the first design, state reasonable assumptions for the rest, and validate them while drawing. Excessive questioning can prevent you from demonstrating a working solution.
What the answer must demonstrate: Connect each clarification to an architectural consequence.
What is a correctness invariant? Give one for an upload-and-download API.
Reveal a model answer
“An invariant is a condition the system must preserve. Here, an incomplete upload must never become downloadable. I represent upload state explicitly and allow downloads only after completion is verified. The download handler can enforce this rule by checking the stored upload state before serving the file.”
Interviewer follow-up
Is “the service is fast” an invariant?
Reveal the follow-up answer
“Fast” needs a performance target and measurement window. The upload rule applies to every relevant request: verify the unchanging object version, then atomically mark it ready only if its metadata still permits that change. A completion arriving after deletion must leave the file deleted. The file store and metadata database do not share one transaction.
What the answer must demonstrate: Give an enforceable rule, not an adjective.
Why start with one application and database?
Reveal a model answer
“It makes the complete request and stored state understandable. I can show which changes must succeed together, verify the rules that keep the data correct, and measure capacity. I split or replicate components when a workload, reliability requirement, or ownership boundary creates a reason, rather than assuming that a distributed diagram is inherently better.”
Interviewer follow-up
What if the interviewer immediately requires global scale?
Reveal the follow-up answer
I still explain the logical operation, then show regional routing, data ownership, and replication. Starting from a clear operation does not require deploying only one machine.
What the answer must demonstrate: Logical clarity should survive changes in physical scale.
A file service averages 2.3 downloads/s but may peak at 1,000/s. Is average QPS enough to choose one server?
Reveal a model answer
“Not from that average alone. I need peak request rate, average and large-file sizes, connection duration, and a per-server load test at the target latency. Traffic can be concentrated into short bursts. I would state the peak assumption and size for it, including a server failure.”
Interviewer follow-up
If metadata QPS is modest but downloads total 12 TB/month, what motivates separate object storage?
Reveal the follow-up answer
Byte delivery dominates metadata traffic: the example has 12 TB/month of downloads. Separating bulk bytes gives an independent delivery path even if metadata QPS is modest.
What the answer must demonstrate: Do not equate average QPS with capacity.
Why define a data model before naming a database product?
Reveal a model answer
“The model tells me what must be stored together and which queries must be efficient. For file sharing I need ownership, upload state, expiry and the public-token mapping checked by identifier. A transactional metadata database can enforce those relationships; I evaluate products after deciding durability, throughput and failure requirements.”
Interviewer follow-up
When might that choice change?
Reveal the follow-up answer
Measured limits or global write requirements might justify partitioning or distributed storage. I would describe the new operational and consistency costs alongside the change.
What the answer must demonstrate: Explain access patterns and constraints.
An upload completion commits but its response is lost. How should a retry behave?
Reveal a model answer
“A timeout means the client does not know the outcome. I keep a stable upload identifier and make the completion operation inspect its existing state. Retrying completion for an already-ready upload returns the same worksheet. I recover the existing outcome before creating a new upload.”
Interviewer follow-up
How would you prove the fix works?
Reveal the follow-up answer
Simulate losing the response after the state transition commits, retry with the same identifier, and verify that exactly one logical worksheet is ready.
What the answer must demonstrate: Explain what happens if the server saves the result but the response is lost.
How do you answer “Why a CDN?” without a buzzword list?
Reveal a model answer
“Repeated downloads request identical bytes. A CDN can reduce origin traffic and serve a nearby copy. I would use versioned public objects where possible. If a link is private or revocable, I must define the authorization and cache lifetime so an old edge copy cannot bypass the promised access policy.”
Interviewer follow-up
Does a CDN remove the metadata service?
Reveal the follow-up answer
No. Metadata still owns upload state, ownership, and permission decisions. Depending on the access scheme, the edge may enforce a limited authorization token or request validation.
What the answer must demonstrate: Name the benefit and the access-policy cost.
How do you close the interview?
Reveal a model answer
“I would recap the agreed user actions, trace the main path briefly, and state the key choices: durable upload states, independent byte delivery, and retryable completion. Then I would identify the first measured scaling limit and one remaining risk, such as revocation latency, with how I would test it.”
Interviewer follow-up
What if the design is unfinished?
Reveal the follow-up answer
Identify the part of the design you have not resolved and explain the next concrete decision you would make. A coherent partial design with clear invariants is better evidence of understanding than pretending every problem is solved.
What the answer must demonstrate: Summarize decisions and limits rather than reciting components.
Does growing traffic mean a modular monolith must become microservices?
Reveal a model answer
“No. I can run multiple instances of the same application when its durable state is shared appropriately. I would extract a capability when independent capacity, releases or ownership justify the extra coordination. In checkout, keeping orders and stock reservations together preserves a useful local transaction; catalog search can scale separately as a derived view.”
Interviewer follow-up
What changes if inventory becomes a separate service?
Reveal the follow-up answer
Order creation and inventory reservation no longer share the original database transaction. The workflow must record progress, identify retries and recover failed or uncertain steps. A reservation needs explicit expiry and confirmation rules. An API call alone does not make the two services commit together.
What the answer must demonstrate: Distinguish server count, deployment boundaries and transaction boundaries.
Blank-page exercise · 45 minutes
Build the answer yourself
Design worksheet sharing from a blank page. Specify upload state, permission checks, and recovery after a lost completion response.
- State four functional requirements and one enforceable invariant.
- Show calculations with units and one peak assumption.
- Trace upload, download, and response-loss retry.
- Defend one choice and explain its cost.
- Say which application boxes are modules versus independently deployed services, and justify one boundary.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
System design interview frameworkHow do you justify each component in the diagram?Recall first, then reveal
A responsibility in a user request or a specific capacity/recovery requirement.
Follow one request.
Return to lessonSystem design interview frameworkWhat makes a tradeoff answer complete?Recall first, then reveal
State the choice, the reason, the cost, and how that cost is handled.
Choice → reason → cost → mechanism.
Return to lessonSystem design interview frameworkWhat does a timeout tell the caller?Recall first, then reveal
The result is unknown; the operation may already have committed.
Unknown is not failed.
Return to lessonSystem design interview frameworkWhen should part of an application become a separate service?Recall first, then reveal
When independent scaling, releases or team ownership justify the extra network failures, API compatibility work and data coordination.
Separate for a benefit; explain the added cost.
Return to lessonFinal revision
Summary and interview notes
A system-design answer turns requirements into a working request path, a stored data model and explicit success and failure rules. Begin with a correct baseline, then justify each change using a workload limit or a required guarantee.
Remember these points
- Separate user actions, quality targets and invariants before choosing components.
- Say when data is safely saved and which saved result a retry returns.
- Verify the immutable stored version, then atomically mark it ready only if current metadata permits publication.
- Scale the measured bottleneck; modest metadata QPS does not imply modest file-delivery bandwidth.
- A monolith can run on multiple servers; extracting services changes deployment and coordination boundaries.
Interview tips
- For every new box, state its responsibility, benefit, added cost and failure behavior.
- Trace a lost response and two concurrent operations; these reveal gaps that a box diagram conceals.
- Finish with the remaining limitation and the next measurement, not a list of product names.
Important qualifications
- The 45-minute allocation is a practice aid, not an employer-wide interview format.
- For immediate revocation, specify exactly when and where a download is authorized. State separately whether already-authorized transfers may finish; revocation cannot remove bytes a user has already downloaded.
Technical references
- Amazon senior SDE interview preparationPrimary employer guidance on practical system-design reasoning and clarifying questions.
- HTTP semanticsReference for request methods, response status, and resource semantics.
- Microservice trade-offsDeployment and module boundaries, benefits and the costs of remote coordination.
- Microservices architecture styleIndependent deployment, data ownership, consistency and operational requirements.
Practice marks stay in this browser.