System designby Learnastra

System-design interview · Extended interviews

Design a notification service

By Anup Rai

Accept one notification intent, plan eligible channel deliveries and recover provider outcomes while respecting current preferences and protecting urgent traffic.

You will learn to

  • Distinguish a business intent, a channel delivery and a provider attempt.
  • Explain opt-out ordering and unknown external outcomes without claiming universal exactly-once delivery.
  • Scale due work, provider quotas and backlog recovery with visible per-channel status.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Message queues, event logs, delivery guarantees, and backpressure · Databases, data models, and ACID transactions · Production readiness: SLI, SLO, observability, and recovery

Workload and timing examples are interview assumptions.

01Define what reliable notification means

Build a shared service for email, push, SMS and an in-app inbox. Trusted product services submit approved templates and business event identities. Audience selection for bulk campaigns, operating an email network and mobile push transport itself are outside scope. A campaign service can expand an audience into individual recipient intents subject to admission limits.

An order service submits event ship-o81 for user U7. The resulting intent N44 plans email delivery D-email and in-app delivery D-app; SMS may be suppressed by preference. The intent is the business request, a delivery is one channel/destination outcome, and an attempt is one provider call. Keeping those identities distinct makes partial success and retries understandable.

A successful API response means the intent is durable. Provider acceptance means a downstream service accepted responsibility; it does not prove device delivery or human reading. Expose those facts separately rather than returning one misleading delivered flag.

Target provider acceptance within thirty seconds for 99% of eligible, immediately due transactional deliveries, measured from durable intent acceptance; marketing can wait. State the denominator: eligible, immediately due messages. Suppressed or scheduled work has a different outcome, and unknown or failed provider calls count as missed delivery targets rather than disappearing from the metric.

Ask which channels and urgency classes matter, what “delivered” should mean, and how each category should handle an uncertain provider result. This design accepts durable intents, exposes channel-specific outcomes and prioritizes eligible transactional messages; it does not promise that a human reads them.

02Functional requirements

  1. Submit notification intents. Accept an authenticated business event, recipient and approved template, returning one durable intent for matching retries.

  2. Plan and schedule deliveries. Choose email, SMS, push or in-app delivery using recipient preferences, verified destinations, due time and quiet-hour policy.

  3. Send and expose outcomes. Create an in-app item or invoke the external provider, then report suppression, pending, accepted, delivered, failed or unknown separately per channel.

  4. Manage and recover delivery. Support preference changes, safe retries, verified callbacks, status lookup and inspection of exhausted or uncertain work.

03Non-functional requirements

These are illustrative interview assumptions, not product facts or measured benchmarks. Confirm them before choosing components, then validate the completed design under the stated workload. Here p95 means the 95th-percentile latency: 95% of measured requests take no longer than that value. Report errors and rejected work alongside latency; a fast failure is not a successful outcome.

  1. Workload and quotas. Plan for 100 million external provider attempts/day and about 11,600 attempts/s at a tenfold burst. Retries consume the same provider quotas; these rates are not assumed provider entitlements.

  2. Urgent-message timeliness. Target provider acceptance within thirty seconds for 99% of eligible, immediately due transactional deliveries. Measure from durable intent acceptance; unknown or failed provider calls count as misses. This is a service objective, not a guarantee of provider uptime or device delivery.

  3. Durability and recovery. Accepted intents, delivery state and incoming provider evidence survive worker/API restarts through the transactional database. Preserve pending work through downstream outages, bound admission before storage fills and verify the database’s deployment failure policy separately.

  4. Consent and privacy. A send authorization after a committed opt-out must refuse the disallowed delivery. A previously authorized external call may already be in flight; stop new authorization if current permission cannot be checked.

  5. Duplicate and outcome semantics. Enforce one local inbox item per delivery. External retry safety depends on provider idempotency/lookup scope; preserve unknown outcomes and explicitly choose duplicate-versus-missing risk when neither is available.

04Complete the asynchronous flow with a database worker

Start with an authenticated API, a relational database and a worker polling indexed due-delivery rows. The order service reliably publishes ship-o81 from its own transaction using an outbox or equivalent retry mechanism. Our notification API cannot atomically commit with the separate order database, so repeated submission must be safe.

The API stores N44 and a planning task. The planner loads the approved template version and recipient preferences, creates stable channel delivery rows and records suppression reasons. D-app is completed by inserting a unique inbox item and updating its delivery record in one local transaction after checking current permission.

For D-email, the worker transaction checks current consent and destination, records authorization to send and releases its locks. It calls the provider using the delivery's stable external key where supported. It records provider acceptance or an unknown outcome, then later applies verified receipts. U7 can continue using the order page while this runs.

A worker restart finds the same durable rows. No broker is required to make this first version asynchronous or recoverable. Even one worker can crash after the provider sends a message but before the worker saves the reply; the local database cannot roll back that send.

Design diagramDurable intent, asynchronous delivery

The API acknowledges stored intent; channel workers later establish delivery facts.

Durable intent, asynchronous deliveryThe API acknowledges stored intent; channel workers later establish delivery facts. product to api: ship-o81, recipient, template; api to db: Store intent and plan; worker to db: Authorize current due delivery; worker to provider: Stable delivery identity; worker to inapp: Unique local itemship-o81, recipient, templateStore intent andplanAuthorize current due deliveryStable delivery identityUnique local itemACTOROrder serviceSERVICENotification APISTOREIntent, preferences,deliveriesWORKERDue-delivery workerEXTERNALEmail / push / SMSproviderSTOREIn-app inboxsync
Read each connection in order
  1. syncship-o81, recipient, templateOrder service → Notification API
  2. syncStore intent and planNotification API → Intent, preferences, deliveries
  3. syncAuthorize current due deliveryDue-delivery worker → Intent, preferences, deliveries
  4. syncStable delivery identityDue-delivery worker → Email / push / SMS provider
  5. syncUnique local itemDue-delivery worker → In-app inbox

05Count intents, deliveries and attempts separately

Assume 100 million provider attempts/day, including retries. That averages about 1,157 attempts/s; a tenfold burst is approximately 11,600/s. If the mix averages 1.1 attempts per external delivery and two external deliveries per intent, it represents about 90.9 million deliveries and 45.5 million intents/day. In-app insertions add database work without provider calls.

At 1 KB of attempt metadata, thirty days require roughly 3 TB before indexes and replicas. Rendered message bodies and destinations may be more sensitive than diagnostic metadata; retain only what is needed rather than multiplying full bodies across every attempt.

A 200 ms average provider call and 11,600 attempts/s imply about 2,320 calls in flight if quotas permit. This is a concurrency estimate, not permission to exceed a provider's contracted rate. A single sequential worker would manage only about five calls/s at that latency.

If a provider receiving 500 attempts/s fails for ten minutes, 300,000 attempts accumulate. After recovery, 1,000/s capacity minus 500/s new traffic leaves 500/s to drain backlog, requiring another ten minutes. Recovery capacity must exceed new arrivals; simply restoring normal throughput leaves the queue permanently behind.

06Persist the plan and the evidence

Interfaces

Request or message Contract
POST /notifications with event, recipient, category and template version Returns accepted N44; duplicate matching events return the same intent.
GET /notifications/N44 Reports queued, suppressed, accepted, delivered, failed or unknown per channel.
PUT /users/me/preferences Changes versioned channel/category policy.

Submit a notification intent

POST /notifications

Notification request values

event:            ship-o81
recipient:        U7
category:         transactional order update
template version: the approved immutable version selected by the caller

Separate channel outcomes for N44

email / D-email: pending at the provider
in-app / D-app:  committed locally
SMS:            suppressed by preference

Stored records

Record Fields or identity Purpose
Intent tenant, event, user, category, payloadHash Stable business meaning and deduplication scope.
Delivery intent, channel, destinationVersion, state, dueAt One logical recipient endpoint across attempts.
Attempt delivery, number, providerKey, outcome Diagnostic record of each external invocation.
In-app item unique by delivery ID Delivery ID (unique) A locally enforceable no-duplicate inbox insertion.

Templates are immutable versions with typed parameters. Callers may select authorized templates, not inject arbitrary destinations or executable markup. Reusing a business identity with different canonical parameters conflicts. Obtain verified email, phone and device registrations from trusted recipient records.

Before the first authorized send, freeze the exact destination and rendered parameters for that delivery. Changing an email address during recovery must not send different content to a new destination under the same provider key. A replacement is a deliberate new logical delivery with a policy for the older uncertain one.

Store incoming provider events in a durable inbox before acknowledging them; store outgoing dispatch work in an outbox with delivery changes. These records make both handoffs retryable after a crash.

08Recover uncertainty without inventing another message

Suppose the email provider accepts D-email but its response is lost. The worker cannot tell whether nothing happened or whether the email is already on its way. Record unknown rather than definitive failure. If the provider supports idempotency, retry the same logical delivery key within its documented scope and retention. If it supports reliable lookup by client reference, query the original operation.

Attempt numbers identify individual calls for diagnosis. Using a new attempt number as a new provider deduplication key would allow every retry to create another message. A worker lease can prevent ordinary concurrent work but cannot stop a paused old process from later calling a provider that does not enforce our lease token.

When the provider has neither idempotency nor reliable lookup, an automatic exactly-once guarantee is unavailable. Choose a category-specific policy: a transactional update may tolerate a documented duplicate risk; marketing may prefer withholding another attempt until reviewed. Keep the uncertainty visible.

Switching providers does not solve this ambiguity because the second provider cannot deduplicate an effect at the first. Fail over after a known failure, or explicitly accept duplicate risk. Definitive invalid-address errors should stop or disable the matching destination version; temporary rate limits use bounded backoff and provider retry guidance.

Request traceProvider acceptance with a lost reply

The second attempt recovers D-email; it does not invent another logical message.

Provider acceptance with a lost replyThe second attempt recovers D-email; it does not invent another logical message. worker to db: Authorize and record D-email; worker to provider: Send with stable delivery key; provider to worker: Acceptance reply lost; worker to db: Record unknown outcome; worker to provider: Lookup or same-key recovery; worker to db: Apply verified original outcomePARTICIPANTEmail workerPARTICIPANTDelivery authorityPARTICIPANTProvider1. Authorize and record D-email2. Send with stable delivery key3. Acceptance reply lost4. Record unknown outcome5. Lookup or same-key recovery6. Apply verified original outcomesyncblocked
Read each connection in order
  1. syncAuthorize and record D-emailEmail worker → Delivery authority
  2. syncSend with stable delivery keyEmail worker → Provider
  3. blockedAcceptance reply lostProvider → Email worker
  4. syncRecord unknown outcomeEmail worker → Delivery authority
  5. syncLookup or same-key recoveryEmail worker → Provider
  6. syncApply verified original outcomeEmail worker → Delivery authority

09Schedule by provider capacity and recipient ownership

Add a broker when measurements show that buffering and separately scaled workers improve delivery of due work. Commit delivery rows and outbox records together; a relay publishes stable delivery IDs. Duplicate publication is acceptable because the worker rechecks the authoritative delivery state before acting. The broker schedules work and does not become the sole copy of consent or message meaning.

Group execution by provider/channel quota and urgency. Reserve capacity for transactional updates so a large campaign cannot occupy every provider slot. Use weighted scheduling and per-tenant limits rather than unlimited strict priority, which may starve lower-priority work. Unused reserved capacity can be borrowed under an explicit rule without removing the transactional reserve during a surge.

Recipient data partitions serve a different purpose. Keep a recipient's preferences, delivery decisions and inbox together so authorization remains local. Queue messages carry enough routing information to reach that owner. Before a new database owner authorizes sends, it must have the current preferences and complete state of deliveries already in flight.

A queue cannot guarantee a thirty-second deadline when backlog exceeds downstream capacity. Reject impossible new work before promising acceptance, defer marketing, and expose oldest eligible age. Adding workers helps only until provider quotas, reputation rules or account-specific limits become the bottleneck.

Design diagramDurable delivery decisions with bounded channel workers

A product submits one stable notification intent. Planning and outbox work create per-channel deliveries; the broker carries delivery IDs to quota-controlled workers. Each worker reloads current policy and delivery state before a provider call. In-app items stay in the recipient’s database, and verified receipts update the same delivery history.

Durable delivery decisions with bounded channel workersA product submits one stable notification intent. Planning and outbox work create per-channel deliveries; the broker carries delivery IDs to quota-controlled workers. Each worker reloads current policy and delivery state before a provider call. In-app items stay in the recipient’s database, and verified receipts update the same delivery history. products to api: Stable intent ID; api to db: Persist intent and plan; planner to db: Plan and read outbox; planner to broker: Publish delivery IDs; broker to workers: Schedule delivery attempts; workers to db: Authorize / record outcome; workers to providers: Send or recover outcome; providers to api: Verified delivery receiptsStable intent IDPersist intent and planPlan and read outboxPublish delivery IDsSchedule delivery attemptsAuthorize / record outcomeSend or recover outcomeVerified delivery receiptsACTORProduct servicesSERVICENotification APISTORERecipient policy /delivery DBWORKERPlanner and outboxrelaySTOREChannel-readyqueuesWORKERChannel worker poolsEXTERNALEmail / push / SMSproviderssyncasync
Read each connection in order
  1. syncStable intent IDProduct services → Notification API
  2. syncPersist intent and planNotification API → Recipient policy / delivery DB
  3. syncPlan and read outboxPlanner and outbox relay → Recipient policy / delivery DB
  4. asyncPublish delivery IDsPlanner and outbox relay → Channel-ready queues
  5. asyncSchedule delivery attemptsChannel-ready queues → Channel worker pools
  6. syncAuthorize / record outcomeChannel worker pools → Recipient policy / delivery DB
  7. syncSend or recover outcomeChannel worker pools → Email / push / SMS providers
  8. asyncVerified delivery receiptsEmail / push / SMS providers → Notification API

10Handle schedules, devices and receipts deliberately

Quiet hours depend on the recipient's named time zone, not the worker's local clock. Store the resolved scheduled instant and the rule version. Define what happens to a local reminder time skipped or repeated by daylight-saving changes, and what a subsequent user time-zone change affects. Marketing cannot bypass opt-out merely because it entered an urgent queue.

Push delivery targets app installations through provider registration tokens. A user may have several devices, and tokens can rotate. Bind registrations to authenticated users, version each destination and disable only the version associated with a definitive invalid-token response. A delayed rejection for an old token must not disable the replacement.

Verified callbacks can repeat or arrive out of order. Deduplicate provider event IDs where supplied and apply channel-specific facts without moving a delivered message backward to queued. Match the provider account and message reference to the intended delivery. An email bounce and an SMS device receipt have different meanings and should not be flattened into one generic success flag.

Inbox reads authorize the recipient and paginate by creation time plus stable item ID. A cursor orders results but is not automatically a frozen snapshot. Notification previews should not reveal content the user is no longer authorized to read through the underlying application.

11Operate recovery and consent as first-class behavior

Monitor intent admission, oldest eligible delivery age, unknown-outcome age, suppression reasons, provider rate-limit responses, bounce rates and retry amplification. Separate channels, tenants and urgency lanes so a healthy aggregate does not conceal an urgent-message backlog. Alert through an independent channel; this service cannot reliably report its own outage through itself.

Dead-letter handling means recording exhausted or invalid work for inspection, not deleting its history. Replaying a delivery preserves its logical identity and frozen parameters. A deployment rollback must not turn already accepted emails into new sends with new IDs merely because local status appears incomplete.

Test opt-out before and after authorization, provider acceptance followed by worker crash, repeated callbacks, destination rotation, daylight-saving scheduling and a ten-minute provider outage. Validate actual drain capacity, not just whether workers restart. Recovery queries also consume provider quota and need a budget.

The principal cost is attempts per channel under provider pricing, plus retained metadata and recovery work. Reduce wasteful retries only when the desired delivery outcome remains supported. Keep personal destinations and rendered bodies out of metric labels, broadly accessible logs and queue names; restrict template editing separately from sending permission.

12Check the design against its requirements

Use the numbered requirements to check the final design. FR refers to the functional list; NFR refers to the non-functional list. Performance rows specify tests still required, not achieved benchmark results.

Requirement Design mechanism Verification and remaining limit
FR1–2; NFR3: accept and plan once Durable intent identity, versioned templates, planning rows and outgoing work. Repeat ship-o81, restart the planner and change a destination during recovery. Recover the same logical intent and frozen delivery parameters.
FR2,4; NFR4: honor current consent Recipient-owned preference and delivery checks serialize opt-out against send authorization. Race opt-out with authorization and make the preference authority unavailable. Refuse later unauthorized sends; do not claim already-sent messages can be recalled.
FR3–4; NFR5: report honest outcomes Stable delivery identity, provider-supported recovery, verified callbacks and unique in-app insertion. Lose a provider response and replay callbacks. Distinguish accepted from read, and never silently turn uncertainty into a new logical send.
NFR1–3: meet urgency under load Reserved transactional capacity, provider quota limits, bounded admission and separate recovery budget. Measure thirty-second success rate during a campaign and a ten-minute provider outage. The outage creates misses; validate net backlog-drain capacity instead of concealing them.

13Rapid revision

Remember: Saved intent, provider acceptance and human reading are different outcomes. Recover the same delivery when a provider reply is lost.

Concern Complete mechanism
Admission Check the sender and permitted template; save one notification per matching business request.
Planning Create a delivery ID for each eligible channel; record why other channels were skipped.
Send permission In a short transaction, recheck current recipient preferences and the allowed destination.
In-app completion Atomically insert one item per delivery and mark it complete.
External completion Reuse the provider key or look up the delivery; keep its status unknown until evidence resolves it.
Retries Delay retries after temporary failures, stop invalid destinations, and keep records that recognize repeated attempts.
Scheduling Honor named time zones, quiet hours and provider quotas; reserve urgent capacity.
Callbacks Verify, persist, deduplicate and apply each channel’s status rules so an old callback cannot overwrite a later known outcome.
Recovery Process faster than new work arrives; inspect old unresolved deliveries as well as queue length.

Close with ship-o81: one accepted intent, email pending at the provider, in-app committed and SMS suppressed. That single example demonstrates the data model, asynchronous boundary, consent policy and honest partial status without requiring a universal delivered boolean.

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

Why separate intent, delivery and attempt?

Reveal a model answer

One business event can create several channel deliveries, and each delivery may require several calls. Separate identities preserve partial status and safe retries.

What the answer must demonstrate: Keep business intent, channel delivery and transport attempt identities separate.

Foundation · Question 2

Does provider acceptance prove the user read the message?

Reveal a model answer

No. Admission, provider acceptance, device delivery and human reading are distinct facts with channel-specific evidence.

What the answer must demonstrate: Distinguish service acceptance, provider acceptance, device delivery and human reading.

Applied · Question 4

The email call times out after possible acceptance. How do you retry?

Reveal a model answer

Recover the same delivery through supported idempotency or lookup. Without either, keep unknown and apply the category’s explicit duplicate-risk policy.

What the answer must demonstrate: Use original provider identity or lookup and state the unsupported-provider limit.

Foundation · Question 5

Why can in-app notifications have a stronger uniqueness boundary?

Reveal a model answer

The item insertion and completed delivery state can share one local transaction with a unique delivery ID.

What the answer must demonstrate: Identify the local transaction boundary that makes inbox insertion unique.

Applied · Question 6

A campaign fills the queue. How do urgent order updates meet their target?

Reveal a model answer

Reserve provider capacity for transactional work and apply tenant fairness. More workers cannot exceed the same provider quota.

What the answer must demonstrate: Reserve urgency within actual provider quotas and calculate net backlog drain.

Follow-up · Question 7

A rejection arrives for an old push token after refresh. What changes?

Reveal a model answer

Disable only the matching destination version, preserving the newer registration.

What the answer must demonstrate: Guard device invalidation by destination version and freeze uncertain send parameters.

Follow-up · Question 8

How should duplicate or reordered callbacks behave?

Reveal a model answer

Verify and persist them, deduplicate available event identities and apply channel-specific facts without regressing known outcomes.

What the answer must demonstrate: Persist verified provider facts and apply channel-specific idempotent transitions.

Blank-page exercise · 45 minutes

Build the answer yourself

Design shipment notifications over email and in-app, then race an opt-out with dispatch and lose the email provider’s response.

  • Agree numbered functional and non-functional requirements, including channel outcomes, urgency and opt-out/duplicate behavior. Then distinguish a business intent, a channel delivery and a provider attempt.
  • Explain opt-out ordering and unknown external outcomes without claiming universal exactly-once delivery.
  • Scale due work, provider quotas and backlog recovery with visible per-channel status.
  • Trace a timeout and a concurrent request using the actual durable records.
  • Review the final architecture against every numbered requirement, including measured bottlenecks, targets still needing validation and remaining failure limits.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Design a notification serviceN44 has a saved in-app item, but its email call timed out. What can status honestly report?Recall first, then reveal

The intent is durable, the in-app item is complete and email is unknown until provider evidence resolves it. Neither proves the user read the message.

Saved intent, provider acceptance and user reading need different evidence.

Return to lesson
Design a notification serviceDoes a queued notification still need permission before sending?Recall first, then reveal

Yes. The queue schedules work; the database checks current consent. The provider’s retry rules govern repeated external calls.

Queue work; authorize sends.

Return to lesson
Design a notification serviceWhen does a delivery backlog shrink?Recall first, then reveal

Only when completion capacity exceeds the rate of new deliveries arriving.

New work uses capacity too; only the spare part drains backlog.

Return to lesson

Final revision

Summary and interview notes

Save each requested notification once, create eligible channel deliveries, and recover uncertain provider results. Check current preferences before sending and keep capacity available for urgent messages.

Remember these points

  • Separate business, channel and attempt identities.
  • Check current consent in the transaction that authorizes sending.
  • Keep a timed-out provider call unknown until evidence resolves it.
  • Protect urgent traffic with actual provider capacity.

Interview tips

  • Name the outcome measured by the delivery target.
  • Trace ship-o81 through email, in-app and SMS, then race an opt-out with send authorization.

Important qualifications

  • Traffic and latency figures are interview assumptions, not claims about a named company's deployment.

Continue after the core interview

Explore the advanced version

The advanced lesson keeps the full detailed design. Use these sections when you want to examine the stronger requirements and failure cases.

Technical references

Practice marks stay in this browser.