A path through the material
Learn, explain, then design.
Each stage builds ideas used by later lessons. Follow the path in order or use prerequisite links to repair a specific gap.
How the guide is organized
Concept chapters begin with a definition and visual model, then explain the mechanism, guarantees, trade-offs, and interview questions. Worked examples illustrate specific operations and failure cases. Each system design opens in the interview version: scope and clarifying questions, separate numbered functional and non-functional requirements, a working baseline, useful estimates, APIs and records, justified scaling, and critical failures. A closing table checks the design against its requirements before the separate rapid-revision table. Its Q&A and exercise use a 45-minute scope. The Advanced link at the top opens the preserved detailed lesson with stronger guarantees and deeper failure analysis.
These stages follow common interview frameworks, including Hello Interview’s delivery framework. Use estimates when they inform a decision and adapt the depth and sequence to the interviewer.
How to practise a lesson
- Learn the definition and mechanism. Then explain a worked example without looking.
- Answer the chapter’s questions aloud; include the reason behind each choice.
- Try the follow-up before revealing its answer.
- Do the timed exercise on a blank page. Write the functional and non-functional requirements first; finish by checking your design against them. Revisit the cards later.
1. Understand one request
Explain what the user can do, what counts as success, which data is needed, and how much traffic to expect.
2. Store and find the data
Trace a query and a write before distributing them.
- Databases, data models, and ACID transactions
- Database indexes: B-trees, composite keys and query access
- Storage engines and data models
- Load balancing: definition, algorithms and failover
- Caching: cache hits, misses, write policies and invalidation
- Proxies: forward proxy, reverse proxy and API gateway
- Data partitioning and sharding
- Consistent hashing and virtual nodes
3. Reason about concurrency and failure
Draw timelines for stale reads, concurrent changes, lost responses, and worker crashes.
- Replication and durability
- CAP theorem: consistency, availability, and partition tolerance
- Consistency models
- Transaction isolation
- Quorums, consensus, leases, and fencing
- Idempotency, retries, and timeouts
- Message queues, event logs, delivery guarantees, and backpressure
- Distributed transactions and sagas
4. Build production and modern foundations
Explain delivery, approximation, permissions, recovery, and observability.
5. Practise complete systems
Rotate through different workloads. Then choose the remaining designs from the library.
- Design a URL shortener
- Design a chat messaging service
- Design a ticket-booking service
- Design a file synchronization service
- Design an API rate limiter
- Design a personalized news feed
- Design a payment system and ledger
- Design a permission-aware RAG knowledge assistant
- Design an LLM inference platform
- Design durable agent workflows
Practise adapting a design to a different prompt
A familiar diagram is a starting point. Identify the rule that changes before reusing its components. Choose one prompt for a 45-minute exercise; these variations do not add requirements to the original exercise.
| Variation | What changes | Where to practise |
|---|---|---|
| Hotel reservations | Reserve room-type capacity on every occupied night. Enough capacity on one night does not guarantee the whole stay. | Multi-night inventory example |
| Wallet transfers | Move existing funds between two accounts while preventing concurrent transfers from spending the same balance. | Internal transfer example |
| Service boundaries | Distinguish adding application instances from giving components separate deployments and transactions. | Monolith and service comparison |
For any variation, restate the requirements, update the data model, trace one successful operation and one failure, then draw the resulting architecture. Explain which parts of the previous design still apply.
Specialist prompts need additional preparation
This library covers general product and infrastructure interviews. It does not yet contain full worked designs for the following specialist prompts. Related chapters provide useful building blocks, but do not replace their distinct requirements.
- Hosted email: notifications use an external delivery provider. Running an email service also needs incoming mail, domain routing, durable server-to-server delivery, mailbox synchronization and abuse handling. See ByteByteGo’s email delivery overview.
- Stock exchange: payments record money movements. An exchange must also sequence orders, reserve trading capacity, decide which buy and sell orders match, process cancellations and publish executions. See ByteByteGo’s exchange overview.
- Untrusted code execution: the job scheduler assumes trusted handlers. An online coding judge additionally needs isolated execution, resource and network limits, and safe cleanup of submitted programs.
A 45-minute mock interview
This is a practice allocation, not a universal interview format. Adjust when the interviewer redirects you.
| Assess your answer | Evidence to look for |
|---|---|
| Agreed requirements | You listed the user actions separately from measurable quality targets, stated assumptions and exclusions, and confirmed the important choices. |
| Understandable | You defined the terms and traced actual records. |
| Grounded | You used workload estimates and required queries to justify design choices. |
| Correct under failure | You showed the commit point, retry, and recovery result. |
| Defensible | You explained a cost and a reasonable alternative. |
| Complete | You checked the final design against the agreed requirements and identified targets that still need measurement. |
Why the 2026 additions are here
The core skills remain practical reasoning, accuracy, reliability, and scalability, as described in Amazon’s interview guidance. The additional topics are a curriculum judgment informed by current production engineering, not a measured ranking of interview frequency.
- Permission-aware retrieval and RAG: Microsoft’s multitenant RAG architecture.
- Model serving and constrained GPU capacity: vLLM’s serving documentation.
- Durable agent sessions and tool execution: Anthropic’s 2026 agent infrastructure article.
- Webhooks, feature flags, tenant isolation, object storage, recommendations, and live video broaden the storage, control, delivery, and interactive workloads you can practise.
Each lesson cites primary technical sources, reviewed in September 2026. Product features depend on their documented configuration; illustrative numbers are not current company measurements. Technology names are implementation options with stated trade-offs, not a requirement to select the newest release. Each chapter ends with summary points, interview tips, and qualifications. Dotted concept links open the relevant explanation in a new tab.