System designby Learnastra

Concept lesson · Foundations

Production readiness: SLI, SLO, observability, and recovery

By Anup Rai

Start here

Definition

Production readiness is the ability to operate a service reliably: measure user outcomes, detect failure, limit damage, deploy changes, and recover. An SLI (service-level indicator) is a quantitative measure of service behavior; an SLO (service-level objective) sets its target over a stated window. The error budget is the unreliability that target permits: for example, the allowed number of bad requests or the allowed downtime, using that SLO’s denominator and window.

Why it matters: A healthy process can still serve slow, incorrect, or incomplete results. Operators need measurements of the operations users depend on, such as uploading and viewing a photo, and tested procedures for recovering those operations after a failure.

The visual modelService-level objectives, error budgets, and observability

An SLI measures service behavior, such as the fraction of photos ready on time. An SLO sets a target over a time window. Compare completion and error rates for the new version with the current version before expanding the rollout.

Service-level objectives, error budgets, and observabilityAn SLI measures service behavior, such as the fraction of photos ready on time. An SLO sets a target over a time window. Compare completion and error rates for the new version with the current version before expanding the rollout. For one million photo uploads evaluated in the rolling 30-day window and a 99.9 percent success SLO, at most 1,000 may miss the defined readiness deadline. Each of the ten equal blocks represents 100 allowed misses. P501 missing the deadline consumes one of the 1,000 permitted bad events, shown as one hundredth of the first block; 999 remain if it is the only miss. Metrics reveal the rate, logs identify a specific job, and traces locate time across stages. A canary compares the new version with the old before rollout expands.Target: uploaded photos ready within 60 seconds30-day budget: 1,000,000 x 0.1% = 1,000 missesEach block = 100 allowed misses100100100100100100100100100100P501: 1 miss used999 remainMETRICready within 60 sLOGP501 took 100 sTRACEqueue wait: 95 sCanary: compare completion and error rates firstDefine success and the time window first; uptime alone does not prove photo readiness.
Read the diagram step by step
  1. For one million photo uploads evaluated in the rolling 30-day window and a 99.9 percent success SLO, at most 1,000 may miss the defined readiness deadline. Each of the ten equal blocks represents 100 allowed misses.
  2. P501 missing the deadline consumes one of the 1,000 permitted bad events, shown as one hundredth of the first block; 999 remain if it is the only miss.
  3. Metrics reveal the rate, logs identify a specific job, and traces locate time across stages.
  4. A canary compares the new version with the old before rollout expands.

Worked example

If 99.9% of one million accepted photos must become ready within 60 seconds, at most 1,000 may miss that target. A 200 response at upload time does not prove that background processing met the objective.

Key takeaways

  • Measure whether the requested operation finishes correctly, including any required background processing.
  • Metrics show the trend; logs and traces explain individual failures.
  • A rollback, failover, or restore is complete only after the user-visible result is verified.

You will learn to

  • Define a user-facing success indicator, objective, denominator, and time window.
  • Use metrics, logs, and traces to distinguish a symptom from its cause.
  • Explain a canary rollback and verify recovery without losing accepted work.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Distributed systems: scalability, reliability, availability and efficiency · Message queues, event logs, delivery guarantees, and backpressure

Workload and timing examples are interview assumptions.

01What is production readiness?

Production readiness means being prepared to run the service through ordinary traffic, changes, overload, and failures. First define what users must be able to do and how reliably and quickly the service must respond. Then decide how to measure whether it meets those requirements and how to recover when it fails. Observability is the ability to understand internal behavior from the metrics, logs, and traces that the service produces.

An SLI (service-level indicator) is a quantitative measure of service behavior, such as the fraction of photos ready within 60 seconds. An SLO (service-level objective) is a target for that measurement over a window. An SLA (service-level agreement) is a commitment with agreed consequences, often contractual; it is not simply another name for an internal SLO. An error budget is the amount of failure the SLO permits over its measurement window, such as the number of requests allowed to miss a completion deadline.

Measure whether the service finishes the operation the user requested, including any background processing required before the result is usable. In the example, upload P501 is accepted immediately, spends 95 seconds queued, takes four seconds to render and one second to publish, and becomes ready after 100 seconds. That event misses a 60-second completion threshold despite a successful acceptance response and running API processes.

Check the complete user operation, protect the resources it needs, deploy changes safely and test recovery. A running process is useful evidence, but does not prove the service is fast enough, saves data correctly or enforces access permissions.

02SLI, SLO, SLA, and error budget: definitions and calculation

A service-level indicator, or SLI, is the measured behavior. A service-level objective, or SLO, is its target over a stated window. For completion, the numerator counts eligible photos ready within 60 seconds. The denominator counts eligible accepted photos whose evaluation period has elapsed. A just-accepted photo cannot be labeled late before its 60-second allowance ends.

Concept in focusHow much of the error budget remains?

This bar represents the 1,000 permitted bad events. It does not represent all traffic.

How much of the error budget remains?This bar represents the 1,000 permitted bad events. It does not represent all traffic. Split a 1,000-event error budget into used and remaining portions. A 99.9% target over 1,000,000 eligible events permits 1,000 bad events. 400 bad events use 40% of that budget, leaving 600; the measured good fraction is 99.96%.99.9% SLO over 1,000,000 eligible eventsError budget: 1,000 bad events allowed400 used600 remainingMeasured result: 999,600 / 1,000,000 = 99.96% good events.The bar shows only the error budget, not all one million events.An SLA is a separate agreement; this target alone is not a contract.

Remember: Allowed bad events minus actual bad events gives remaining budget.

Read the diagram
  1. Split a 1,000-event error budget into used and remaining portions.
  2. A 99.9% target over 1,000,000 eligible events permits 1,000 bad events.
  3. 400 bad events use 40% of that budget, leaving 600; the measured good fraction is 99.96%.
Try from memoryHow many additional bad events fit in the current fixed window?

600, assuming the window still contains exactly 1,000,000 eligible events and the target remains 99.9%.

Decision Example metric
Operation Valid uploaded photo becomes viewable
Good event Ready no later than 60 seconds after acceptance
Denominator Eligible accepted photos with an elapsed evaluation period
Target and window At least 99.9% over rolling 30 days
Separate guardrail Valid upload attempts are accepted successfully

Calculate the error budget

For one million evaluated photos, the 0.1% allowance permits at most 1,000 bad completion events. This allowance is an error budget. P501 is one bad completion event that consumes this budget; one late photo alone does not prove the aggregate 30-day 99.9% SLO was violated. Specify whether unsupported file types, canceled uploads, and failures caused by our service count. Exclusions should reflect the contract, not hide inconvenient incidents. Availability, timely completion, and correctness can require different indicators.

Define exactly which events enter the window

Handle no traffic and missing telemetry

When there are zero eligible events, the ratio is undefined, not 100% healthy. Use a no-data signal and the separate acceptance indicator or synthetic check. A time-based 99.9% availability target over 30 days permits 43.2 minutes of bad time, but that is a different denominator from the one-million-photo event budget. Do not convert between them without traffic assumptions.

A synthetic check performs a controlled test operation, such as uploading a test image and verifying that it becomes viewable. It can reveal a broken path when real users are inactive. Report that test separately from the real-user completion ratio rather than using it to invent a denominator for a no-traffic period.

03Observability and the four golden signals

Metrics are numerical measurements over time. For this service, track upload demand, timely completion, queue age, worker capacity, and errors. The classic four signals are latency, traffic, errors, and saturation: how long work takes, how much arrives, what fails, and which resource is nearly full. Google SRE monitoring.

The 100-second completion is the symptom. High queue age tells us where to investigate; it is not yet the cause. CPU may be low because a worker-concurrency setting is too restrictive, not because there is no demand.

Observation What it tells us What it does not prove
Upload responses succeed Acceptance path is responding Photos become ready promptly
Queue age rises Work is waiting longer The queue service is broken
Worker CPU is 25% CPU is not fully occupied Sufficient workers are active
New-release cohort is slower Release is a useful suspect Causation without further inspection

Break down metrics by processing stage and software version. Control label cardinality: the number of distinct label values and combinations that create separate time series. A separate time series for every photo ID would be costly; IDs belong in targeted event records and traces.

04Logs and distributed traces: locate the missing 95 seconds

Logs record individual events; structured fields make those records searchable. Traces connect work across stages so we can follow one request or asynchronous job. A span records one timed operation within a trace, such as a database call or a worker processing a photo. Carry a correlation identifier that links records for the same job without exposing secrets or personal data through acceptance, queue delivery, rendering, and publication. For asynchronous work, preserve the relationship even when it is represented by a trace link rather than one continuously open call.

Concept in focusA trace shows time spent within one request

Bar width is elapsed time. Child spans overlap the parent’s time and must not be added to it.

A trace shows time spent within one requestBar width is elapsed time. Child spans overlap the parent’s time and must not be added to it. Locate the database and remote-call durations inside a 100 ms API span. The database call runs from 10 to 30 ms. The remote call runs from 35 to 90 ms. Other work or waiting occupies the unlabelled intervals.One request: where did its 100 ms go?API100 msDB call20 msRemote call55 ms0 ms25 ms50 ms75 ms100 msThe child bars sit inside the parent span. Gaps are other work or waiting.

Remember: Read the timeline to locate the slow segment.

Read the diagram
  1. Locate the database and remote-call durations inside a 100 ms API span.
  2. The database call runs from 10 to 30 ms. The remote call runs from 35 to 90 ms.
  3. Other work or waiting occupies the unlabelled intervals.
Try from memoryShould the API time be calculated as 100 + 20 + 55 ms?

No. The child spans occur inside the 100 ms parent interval; adding them double-counts their time.

P501 event Elapsed time Evidence
Upload accepted durably 0 seconds Acceptance record
Worker begins 95 seconds Queue/job trace
Rendering finishes 99 seconds Worker span or event
Photo becomes ready 100 seconds Publication record

Rendering took four seconds and publication one. Almost all delay was before work began. We inspect the new worker release and discover that its concurrency limit was unintentionally reduced. That mechanism fits both the queue wait and low CPU.

Logs must not copy private image contents, access tokens, or unnecessary personal data. A photo ID and authorized diagnostic lookup are usually more useful than dumping the entire payload into an unrestricted log.

A practical implementation can instrument request and worker spans with OpenTelemetry, propagate trace context in the job metadata, and export selected traces and structured logs to a backend. Use the durably stored job record to decide whether a job completed; sampled traces are diagnostic evidence, not a complete SLO denominator. Cross-host timestamps may differ, so record stage durations with monotonic timers, which measure elapsed time without jumping when the system clock is adjusted, and account for clock uncertainty when subtracting timestamps from different machines.

Worked example diagramP501 waits 95 seconds, renders for 4, and publishes for 1: 100 seconds total. It misses the 60-second deadline. Metrics detect the symptom; trace and controlled release evidence support the mitigation decision. One miss alone does not establish a 30-day SLO breach.
Production readiness: SLI, SLO, observability, and recovery: architecture diagram1. Upload P501 to 2. API: accepted at 0s: valid authenticated upload; 2. API: accepted at 0s to 3. Queue: wait 95s: stored processing job; 3. Queue: wait 95s to 4. Worker: render 4s: worker begins at 95s; 4. Worker: render 4s to 5. Publish: 1s: render done at 99s; 5. Publish: 1s to 6. Outcome: ready at 100s: ready at 100s; 6. Outcome: ready at 100s to 7. Trace + release evidence → action: completion misses 60s objective1 → 2: valid authenticated upload2 → 3: stored processing job3 → 4: worker begins at 95s4 → 5: render done at 99s5 → 6: ready at 100s6 → 7: completion misses 60s objective01Upload P50102API: accepted at 0s03Queue: wait 95s04Worker: render 4s05Publish: 1s06Outcome: ready at100s07Trace + releaseevidence → action
  1. 1 → 2valid authenticated uploadUpload P501 → API: accepted at 0s
  2. 2 → 3stored processing jobAPI: accepted at 0s → Queue: wait 95s
  3. 3 → 4worker begins at 95sQueue: wait 95s → Worker: render 4s
  4. 4 → 5render done at 99sWorker: render 4s → Publish: 1s
  5. 5 → 6ready at 100sPublish: 1s → Outcome: ready at 100s
  6. 6 → 7completion misses 60s objectiveOutcome: ready at 100s → Trace + release evidence → action

05Actionable alerts and error-budget burn rate

A dashboard helps investigation; an alert asks someone to act. Paging on every brief CPU spike creates noise and does not necessarily protect the completion objective. Tie urgent alerts to significant user-impact or rapid budget consumption, with enough evidence to identify the affected service and likely response.

Suppose a recent window has 2% late photos while the SLO allows 0.1%. The burn rate is 2% / 0.1% = 20: the service is consuming its error allowance at twenty times the reference rate under that measurement. Use both shorter and longer windows so a severe ongoing problem is detected without treating a tiny transient sample as a sustained incident. SLO alerting reference.

Also monitor correctness constraints. A timely response that exposes a private photo is not a successful product outcome. Audit access-control decisions and check that rules such as “only authorized users can view a private photo” hold; latency metrics cannot establish confidentiality. The security-and-multi-tenancy chapter explains where and how to enforce those access checks.

For the illustrative 30-day window, a sustained 20× burn would consume a full window's budget in about 30 / 20 = 1.5 days under steady traffic and the same bad-event definition. That is a planning approximation, not a promise about a rolling window with changing request rates. Each paging alert should identify the affected objective, the team responsible for responding, a link to diagnostic information, and the first safe action to reduce the impact. Route slower budget erosion to a nonurgent work queue rather than paging on every symptom.

06Canary deployments, rollback, and backlog recovery

A canary release sends a limited portion of work to a new version before broad rollout. Compare workers running the new version with a control group running the current version on similar jobs. Measure whether photos become ready on time as well as whether the worker processes are running. In this example, route comparable jobs to a small canary worker pool with its own bounded queue so queue wait can be attributed to that pool. The canary shows elevated waiting and the reduced concurrency setting; stop expansion and restore the known-good configuration. If old and new workers instead pull from one shared queue, queue age is a shared symptom, not a per-version causal measurement. Compare per-version processing throughput and controlled workload evidence before attributing the delay.

Recover work accepted during the rollout

Keep data formats compatible

A schema change may prevent a simple binary rollback if the old code cannot read new data. Deploy changes in stages so that old and new application versions can both read the stored data during the transition. Feature flags can enable a new behavior separately from deploying the code. Limit the blast radius—the number of users or resources affected by one mistake—through gradual deployment and workload isolation. Keep a clear incident record of the symptom, change, action, and measured recovery.

Blue-green deployment: switch environments

A blue-green deployment prepares a second application environment, validates it, then shifts traffic from the old environment to the new one. It gives a clear traffic rollback target, but temporarily duplicates capacity and still needs connection draining: stop sending new work to the old environment while allowing its existing requests or connections to finish. A canary instead exposes a bounded cohort to the new version before broader rollout; either pattern needs comparable outcome measurements.

What traffic rollback cannot undo

07Disaster recovery: RPO, RTO, failover, and restore

A lost worker can be replaced and its jobs redelivered. A lost region may require a wider failover. A replicated bad deletion may require restoring older history. Choose the response from the actual failure, rather than treating every incident as a request to restart machines.

Recovery point objective, RPO, describes the acceptable loss of recent data measured in time. Recovery time objective, RTO, describes the target time to restore useful service. Both require tested procedures and measured results. For accepted photos, verify that the original files, records of pending processing jobs, publication status, and access permissions all survive recovery. Restoring one database does not by itself prove that users can upload and view photos again. Recovery guidance.

The multi-region-and-disaster-recovery chapter develops region placement and failback. Here the operational lesson is evidence: rehearse the recovery, check the customer-visible result, and record whether the objectives were met. A successful backup command or green failover control-plane status is only partial evidence.

Recovery planning also identifies the team responsible for recovery, a runbook with step-by-step instructions, accessible credentials and keys, and the dependencies needed to serve the recovered data. Verify that the remaining system can handle the required load when a server, zone, or region covered by the recovery plan is unavailable, and perform a restore to an isolated environment before relying on the procedure. Recovery point is a target: asynchronous replication lag must be measured to determine whether the observed lost work meets that target. Backups that share the same destructive permissions and retention policy as live data can fail together.

08Interview answer: explain how you know a service is healthy

Interviewer: “How will you know the upload service is healthy?”

Candidate: “I would measure both valid upload acceptance and whether accepted photos become ready within the agreed time. P501 returned success immediately but took 100 seconds, so an HTTP-success dashboard would miss the completion failure.

“I would trace acceptance, queue wait, rendering, and publication. The 95-second wait points toward processing capacity, and the canary’s reduced concurrency setting explains it. I would roll back that setting, verify the backlog drains, and check that replayed jobs preserve one correct result and private access. Alerts would focus on completion failures and error-budget burn.”

This answer connects monitoring to action: what the user needed, which measurements distinguish likely causes, what change is safe to undo, and how to check recovery. A monitoring box in a diagram needs those explanations.

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

Define SLI, SLO, SLA, and error budget with one user-visible example.

Reveal a model answer

An SLI (service-level indicator) is a measured user-visible behavior. An SLO is its target over a window. An SLA is an agreement about service commitments and consequences, often contractual. An error budget is the amount of failure permitted by the SLO.

For a photo service, measure the fraction of eligible accepted photos that become ready within 60 seconds, after each photo's evaluation period has elapsed. Set an illustrative SLO of at least 99.9% over 30 days. Of one million evaluated photos, at most 1,000 can miss the target. A separate SLA could specify a contractual commitment and credits; it need not use the same threshold as the internal SLO. Also measure acceptance so the service cannot make its completion ratio look good by rejecting every upload.

What the answer must demonstrate: A percentage without a denominator and window is incomplete.

Applied · Question 2

Could accepting no uploads make your completion SLO look perfect?

Reveal a model answer

“Yes, if it treats no data as success, or shows only accepted uploads and hides rejected attempts. Zero evaluated photos proves nothing about completion. Also measure how many valid attempts are accepted, signal missing data and use a test upload when useful. Then we can distinguish failure to accept uploads from failure to process them.”

What the answer must demonstrate: Beware metrics that improve by refusing useful work.

Applied · Question 3

For 1,000,000 evaluated operations and a 99.9% success SLO, what is the error allowance? What burn rate does a 2% bad-event rate represent?

Reveal a model answer

“At 99.9%, one million evaluated operations allow one thousand bad events. If a recent window has two percent bad events against a 0.1 percent allowance, its burn rate is twenty. I would interpret that with traffic and window size before deciding how urgently to page.”

What the answer must demonstrate: Keep percentage points and ratios distinct.

Applied · Question 4

A photo takes 100 seconds: 95 queued, 4 rendering, 1 publishing. What does this trace reveal that low CPU usage does not?

Reveal a model answer

“The trace assigns ninety-five seconds to waiting, four to rendering, and one to publication. Low CPU cannot tell me whether concurrency is accidentally restricted or demand is absent. The trace locates the delay; release/configuration evidence then helps identify the cause.”

What the answer must demonstrate: Separate symptom, location, and causal evidence.

Foundation · Question 5

When do you use metrics, logs, and traces?

Reveal a model answer

“Metrics show aggregate trends and support alerts. Structured logs record individual events. Traces or correlated job events connect a specific journey across stages. For P501 I use metrics to detect late completion and the trace plus targeted logs to explain where it waited and which release handled it.”

What the answer must demonstrate: Choose the evidence type according to the question.

Applied · Question 6

What should the canary compare before full deployment?

Reveal a model answer

“Compare similar workloads, user-visible completion, throughput and stage delay, with enough observations to distinguish a signal from noise. For worker changes, separate canary and control pools can make queue-wait attribution meaningful. If both versions share a queue, rising age affects the cohort comparison and cannot by itself blame one version. I would inspect per-version throughput/configuration and stop expansion or revert when the canary violates the agreed guardrails.”

What the answer must demonstrate: Deployment safety includes data compatibility.

Follow-up · Question 7

The old worker version is back. Can you close the incident?

Reveal a model answer

“Only after verifying the backlog drains and timely completion recovers. I also check that retrying a job did not publish duplicate results or repeat other side effects, and that photo permissions remain correct. Restoring the old version is a recovery step; I still need to verify that users can upload and view photos successfully.”

What the answer must demonstrate: Verify recovery under continuing load.

Follow-up · Question 8

How do RPO and RTO change your recovery exercise?

Reveal a model answer

“RPO tells me how much recent accepted work may be lost; RTO tells me how soon useful service should return. I would measure both during a drill and verify original files, records of pending jobs, publication status, and permissions, rather than timing only a database restore command.”

What the answer must demonstrate: Recovery objectives apply to the service outcome.

Blank-page exercise · 18 minutes

Build the answer yourself

Design a dashboard and incident response for P501 becoming ready at 100 seconds despite a successful upload response. Compare a canary worker release with the control.

  • Define eligible requests and separate acceptance from timely completion.
  • Calculate the error allowance and burn-rate example.
  • Use a trace to identify where the 100 seconds was spent.
  • Describe rollback, backlog recovery, and a check that private photos remain private.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Production readiness: SLI, SLO, observability, and recoveryWhat does an SLI measure?Recall first, then reveal

An actual service behavior, such as the fraction of accepted photos ready within 60 seconds. The SLO is the target and evaluation window.

Indicator measures; objective targets.

Return to lesson
Production readiness: SLI, SLO, observability, and recoveryHow many bad events does 99.9% permit among one million evaluated photos?Recall first, then reveal

At most 1,000 for that defined metric and window.

One in a thousand is the allowance.

Return to lesson
Production readiness: SLI, SLO, observability, and recoveryWhy is P501 a bad completion event despite HTTP success?Recall first, then reveal

Acceptance completed, but the photo waited 95 seconds and became ready at 100 seconds, beyond the 60-second good-event threshold. It consumes error budget; the aggregate SLO depends on all evaluated events.

Upload accepted ≠ photo ready.

Return to lesson
Production readiness: SLI, SLO, observability, and recoveryWhen is rollback actually successful?Recall first, then reveal

The known-good configuration is restored, queued work is draining, photos become ready on time, and their contents and access permissions are correct.

Changed back is not yet recovered.

Return to lesson

Final revision

Summary and interview notes

Production readiness means setting measurable reliability targets, limiting how many users a faulty release can affect, and testing recovery procedures. A running process or successful rollback command does not prove recovery: users must again be able to complete their operations, queued work must drain, and data and permissions must remain correct.

Remember these points

  • An SLI (service-level indicator) is a measurement, an SLO is its target and window, and an SLA is an agreement with consequences.
  • A 99.9% event SLO over one million evaluated photos allows 1,000 missed outcomes; zero events supplies no success evidence.
  • Check each upload once when its readiness deadline arrives, including uploads still unfinished. Counting only completed jobs hides stuck work.
  • A 2% bad-event rate against a 0.1% allowance is 20× burn, interpreted with traffic and window size.
  • Metrics identify impact; traces and logs investigate causes; controlled canary evidence supports a release decision.

Interview tips

  • Write the denominator, deadline, exclusions, and rolling-window rule before drawing a dashboard.
  • Separate acceptance, timely completion, correctness, and confidentiality instead of treating an HTTP success response as proof of all four.
  • For a worker canary, ask whether shared queues and workloads make the cohorts comparable.

Important qualifications

  • Sampled traces cannot stand in for a complete SLO event counter; missing telemetry needs detection.
  • A duration budget and an event budget are different measures even when both use 99.9%.
  • RPO and RTO are objectives to demonstrate in a drill, not guarantees created by configuring replication or a backup job.

Technical references

Practice marks stay in this browser.