System-design interview · Extended interviews
Design a metrics, logging and tracing platform
Build an operational telemetry platform that links metrics, logs and traces, bounds ingestion and query costs, and distinguishes missing evidence from healthy behavior.
You will learn to
- Trace an application request from collection to dashboard, alert and diagnostic evidence.
- Explain metric types, cardinality, reset-aware rates and valid percentile aggregation.
- Protect applications and accepted telemetry during overload while making freshness and sampling visible.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Capacity estimation: throughput, latency, concurrency and storage · Database indexes: B-trees, composite keys and query access · Message queues, event logs, delivery guarantees, and backpressure
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Start with an incident the platform must explain
Checkout request req81 becomes slow and times out while calling inventory. The platform should let an operator see increased checkout latency, locate the timeout log and inspect the trace linking checkout to inventory. Metrics describe aggregate behavior, logs describe individual events, and traces connect timed operations called spans across a request. These signals complement one another rather than being three interchangeable storage formats.
Build operational monitoring with bounded queries and explicit loss reporting. Required financial or security audit records need a separately specified zero-loss or durable-admission contract; ordinary debug telemetry must not silently inherit that promise. Include ingestion, dashboards, structured log search, trace lookup, alert rules and retention. Exclude arbitrary joins over every event ever collected.
Target recent accepted metrics becoming visible within thirty seconds for 99% of samples under normal load and target bounded dashboard queries finishing within two seconds p95. Durable ingestion acceptance is a separate milestone from search visibility. A trace may be sampled or incomplete; a missing span does not prove the service was never called.
An empty graph during a collection outage is missing evidence, not a measured zero. The interface and alert engine must show stale or incomplete data instead of declaring checkout healthy.
Ask whether this is operational telemetry or mandatory audit evidence, which incident queries must be fast, and how much collection loss is acceptable during an outage. Choose operational metrics, logs and traces here with reported dropping of telemetry before acceptance, durable central acceptance and bounded recent-data queries.
02Functional requirements
Collect diagnostic signals. Ingest metric samples, structured logs and linked trace spans with tenant identity and useful request correlation.
Investigate incidents. Provide recent dashboards, bounded structured-log search and trace lookup, showing freshness, sampling and missing evidence.
Evaluate alerts. Evaluate versioned rules over suitable data, preserve pending/firing transitions and notify operators of meaningful state changes.
Manage history. Retain and query the selected recent history, support documented archive/downsample policies and expose rejected or lost telemetry.
03Non-functional requirements
These are illustrative interview assumptions, not product facts or measured benchmarks. Confirm them before choosing components, then validate the completed design under the stated workload. Here p95 means the 95th-percentile latency: 95% of measured requests take no longer than that value. Report errors and rejected work alongside latency; a fast failure is not a successful outcome.
Ingestion and retention. Plan for ten million active series sampled every fifteen seconds, about 666,667 samples/s, plus 100,000 logs/s at 500 bytes each. Use thirty days of metric retention and seven days of indexed logs for the estimates; sampling and archive policies are separate choices.
Visibility and query latency. Target recent accepted metrics becoming queryable within thirty seconds for 99% of samples under normal load, and bounded recent dashboard queries completing within two seconds p95. Disclose slower archive queries and measure timeout/rejection rates alongside latency.
Loss and recovery boundary. Before central acceptance, collectors may shed ordinary telemetry only under a bounded, visible policy. After acceptance, preserve retained records through API/writer restarts and replay them; configure and test the chosen backend’s replica failure guarantees explicitly.
Diagnostic correctness. Compute reset-aware rates and compatible histogram aggregation; do not interpret missing samples as zero or incomplete traces as proof an unobserved call never happened.
Isolation and privacy. Enforce tenant, cardinality, byte and query-work limits. Redact sensitive payloads before persistence and prevent historical queries or unbounded collector buffers from exhausting current application/ingestion capacity.
04Complete the path from request to diagnosis
Start with instrumented applications, a local collector, a time-series engine, a structured log store, a trace store and a query/alert service. OpenTelemetry-style instrumentation can attach request and trace IDs consistently. The collector batches records so each application request does not open its own telemetry network connection.
For req81, checkout increments a request counter, records its duration in a histogram and emits a structured error log with the trace ID. The collector submits a bounded batch. The receiving store acknowledges according to its configured durable-write policy; the collector can then release that accepted batch from its retry buffer.
The dashboard reads a recent bounded interval for checkout. An alert detects a sustained error ratio, records its pending/firing transition and links to the evaluated time range. The operator opens the trace, sees time spent waiting for inventory and searches logs by the same request ID. This is a complete useful product before a custom streaming storage system is added.
The baseline includes tenant authorization, redaction and buffer limits. If telemetry collection consumes unlimited memory or blocks application requests indefinitely, a monitoring outage can cause the checkout outage it was intended to diagnose. State from the beginning which records a full collector may drop before service acceptance.
Each signal has a different query purpose; the collector batches all three.
Read each connection in order
- asyncMetrics, logs and linked spansCheckout instrumentation → Bounded collector
- asyncMetric batchesBounded collector → Time-series store
- asyncStructured eventsBounded collector → Structured log store
- asyncTrace/span identitiesBounded collector → Trace store
- syncError and latency queriesDashboard and alert service → Time-series store
- syncRequest investigationDashboard and alert service → Structured log store
- syncDependency timingDashboard and alert service → Trace store
05Define metric identity and aggregation semantics
A time series is a sequence of measurements with one metric name and label set. For example:
One checkout error-counter series
requests_total{
service="checkout",
instance="c3",
status="500"
}
Changing the instance or status label identifies a different series. A counter increases as events occur and may reset after restart. A gauge reports a current value such as queue depth. A histogram records a distribution through counts in value ranges.
| Question | Correct computation |
|---|---|
| How many requests per second? | Compute a reset-aware rate for each original counter series, then sum rates. |
| What is current queue depth? | Read the gauge, respecting its observation time and aggregation meaning. |
| What is service-wide p99 latency? | Merge compatible histogram counts and estimate the quantile of the combined distribution. |
| What happened to req81? | Search structured logs or the trace ID; do not create a metric series for that request. |
If one instance's counter restarts from 1,000 to 5, summing raw counters before handling resets creates misleading drops. Likewise, averaging two instance p99 values does not produce the service p99: the summaries omit the underlying distribution and request volumes. Histogram bucket boundaries determine approximation error, so use compatible boundaries or a backend-supported compatible histogram representation.
06Count series identities as well as payload bytes
Assume ten million active series sampled every fifteen seconds. That produces about 666,667 samples/s. At an illustrative sixteen bytes for a scalar timestamp/value pair, raw payload is 10.67 MB/s or 0.922 TB/day. Thirty days with three copies is about 83 TB before compression, indexes and metadata. Histograms may require multiple bucket series or larger native samples, so their actual encoding must be included separately.
Logs at 100,000 events/s and 500 bytes each produce 50 MB/s or 4.32 TB/day. Seven days is 30.24 TB before replication and indexing. Indexing every arbitrary field can multiply the storage cost; select fields used by the promised incident queries.
Cardinality is the number of distinct series. A family spanning 100 services, twenty regions, ten statuses and fifty instances could create one million series if every combination exists. Adding request IDs or user IDs can create a new series for nearly every request. A byte limit alone does not protect the series dictionary from that churn.
Set per-tenant limits on active series, new series per minute, log bytes and query work. Measure real label distributions and compression rather than treating a theoretical maximum or raw payload estimate as a final machine count.
07Expose identity, time and completeness
Signal records
| Record | Fields carried by the record |
|---|---|
| Metric batch | Metric name, canonical labels, timestamps, values and producer identity. |
| Log event | Stable event ID, service, event time, ingestion time, severity and trace/request IDs. |
| Trace span | Trace ID, span ID, parent relationship, start time, duration and approved attributes. |
Example structured error-log identity and timing
{
"eventId": "log-req81-1",
"service": "checkout",
"eventTime": "2026-10-07T09:00:00.000Z",
"ingestionTime": "2026-10-07T09:00:00.200Z",
"severity": "ERROR",
"requestId": "req81",
"traceId": "trace-81"
}
Ingestion and query contracts
| Interface | Required information |
|---|---|
| Ingest response | Accepted records or explicit per-record rejection, plus the durability boundary. |
| Query | Authenticated tenant, bounded time range, filters, resolution and output limit. |
| Query response | Results plus freshness, completeness and sampling indicators. |
Event time is when the application observed something; ingestion time is when the platform accepted it. Both matter: an old log may have been buffered by its source rather than delayed by central storage. Specify the allowed out-of-order window and route later data to a defined correction or archive path.
Retried batches retain stable record identities. Two legitimate logs with identical timestamps are still different events. For a simple log store, unique tenant/event IDs make repeated insertion harmless; metric sample identity follows the selected backend's documented series/time and conflict rules. Reject conflicting reuse rather than silently counting a repeated record as new.
Responses must make partial acceptance explicit so the collector retries only the appropriate records. Bound the deduplication/retry horizon and account for its metadata rather than promising infinite remembered identities.
08Add a durable buffer and isolate query resources
If storage outages or bursts exceed direct-ingestion capacity, insert a replicated ingestion log between authenticated gateways and the signal backends. Gateways validate tenants and limits, then acknowledge only after the selected durable append policy. Writers consume accepted records into the time-series, log and trace stores. Search visibility can now lag while accepted data remains buffered.
Use established backends and their supported ingestion/replay integration rather than designing a new chunk-publication protocol during the first interview. Where duplicate-free storage is promised, backend writes must recognize stable record IDs when writers retry. If a backend can expose duplicates, disclose that behavior and deduplicate the relevant query rather than claiming the broker alone supplies exactly-once storage.
Partition metrics by tenant and series identity, logs by tenant/time plus a distribution key, and trace lookup by trace ID. A large tenant needs several partitions; one tenant label must not force all its work onto one machine. Within each backend partition, preserve its required record order and allowed lateness.
Separate expensive historical queries from current ingestion through dedicated resource pools and scanned-byte/concurrency limits. A user investigating an incident should not make every current alert stale by launching an unlimited seven-day search on the same disks.
Acceptance and visibility are separate; backlog age exposes their distance.
Read each connection in order
- syncDurable accepted batchesAuthenticated ingestion → Replicated ingestion log
- asyncReplayable recordsReplicated ingestion log → Signal-specific writers
- asyncSupported idempotent ingestionSignal-specific writers → Metrics / logs / traces
- syncBounded tenant/time queriesBudgeted query workers → Metrics / logs / traces
- syncKnown timestamped signalIndependent freshness probe → Authenticated ingestion
- syncVerify visibility and ageIndependent freshness probe → Budgeted query workers
09Design the questions before indexing everything
For a checkout dashboard, first select series using low-cardinality labels, then read only the requested time interval and resolution. Apply counter resets before aggregation and combine histogram distributions correctly. Return the data-through time and disclose missing partitions. A partial result must not look like a complete service-wide graph.
For req81, narrow by tenant, service and event-time range, then use indexed request or trace IDs. Return ingestion time so the operator can identify delayed evidence. Log pagination needs a stable tie-breaker such as event ID and either a fixed query snapshot or declared live-search behavior. A cursor alone does not freeze incoming records.
A trace view assembles spans under the trace ID and shows missing relationships. Head sampling decides near request start and propagates the decision; tail sampling waits for spans and can favor error traces, but costs buffering and cannot guarantee all late spans arrive before its decision. State sampling rates and incompleteness in the UI.
Cache repeated bounded dashboard queries using tenant, filters, resolution and the relevant freshness boundary. Late data may require invalidating or refreshing affected ranges. Downsampling saves storage but permanently removes time detail; retain count and sum for means and compatible histogram information for distribution queries.
10Make alert state depend on fresh evidence
An alert evaluator might run every fifteen seconds and fire when checkout's error ratio stays above a threshold for a specified duration. Record the rule version, last evaluated boundary, pending start and current state. This prevents a restart from inventing a new alert every evaluation or forgetting all pending progress.
The evaluator first checks whether the input is sufficiently fresh and complete. Missing data produces a stale or unknown state, not a zero-error rate and automatic resolution. A previously firing alert should follow an explicit missing-data policy; silence is not proof that the incident ended.
To fire after a threshold holds continuously, the evaluator needs observations covering the whole interval. If the platform was unable to observe a five-minute interval, retaining an old pending timestamp does not prove the threshold held throughout it. Reconstruct that interval from retained samples if possible; otherwise restart the duration check according to the rule rather than counting unseen time as demonstrated failure.
Persist state transitions and use stable notification identities so retries do not create repeated pages for the same transition. Include the evaluated query and time range in the notification. Monitor the observability platform with an independent small heartbeat or external probe; relying only on its own failed ingestion path can conceal its complete outage.
11Bound buffers and describe the acceptance boundary
| Failure | Response |
|---|---|
| Collector cannot reach ingestion | Buffer to a fixed byte/time limit, retry with backoff, then apply disclosed shedding and count losses. |
| Central durable log cannot accept safely | Stop acknowledging new batches; do not claim collector memory is central durability. |
| Signal writer fails after acceptance | Retain accepted input for replay, exposing delayed query visibility. |
| Query is too broad | Cancel or reject it with a narrower-range suggestion; protect ingestion and alert capacity. |
| Cardinality suddenly explodes | Reject or quarantine invalid dimensions and expose the offending instrumentation. |
The before/after acceptance distinction is essential. Before central acceptance, configured collector buffers can overflow and ordinary telemetry may be lost. After acceptance, the platform has promised retention under the configured backend durability policy, including survival of API/writer restarts. It must not overwrite unprocessed records merely because the buffer is full. Tighten admission before that happens.
At 100,000 log events/s, a one-hour outage accumulates 360 million events. A recovered writer handling 150,000/s while 100,000/s new traffic continues drains at only 50,000/s, taking two more hours. Reserve recovery capacity and measure oldest unprocessed age. Returning to normal ingest throughput alone does not catch up.
12Control cost without destroying incident evidence
Apply tenant-scoped ingestion credentials, label allowlists and redaction before persistence. Private data in labels both leaks information and creates expensive series identities. Restrict retention and alert-route changes, and avoid raw payloads in diagnostic labels about the platform itself.
Monitor accepted-to-visible delay, collector drops, series churn, writer backlog age, query scanned bytes, alert freshness and backend errors. A fast ingestion endpoint can still feed a platform whose dashboards are hours behind. The independent heartbeat should exercise a known write and read, not merely check that a process responds.
Retain recent high-resolution metrics and indexed logs on fast storage, then archive or downsample according to the actual investigative needs. Reducing seven-day indexing can save substantial cost but may make old request searches slower or unavailable. Show those limits rather than promising the same query latency across every retention tier.
Test lost batch acknowledgments, duplicate event IDs, counter resets, incompatible histogram changes, collector overflow and backend replay. During backend upgrades, compare counts, sums and histogram buckets on representative traffic. A custom immutable-block engine would also need to prove that publication and cleanup cannot expose duplicate or missing blocks; selecting a proven backend lets the first interview focus on signal correctness and the complete operating contract.
13Check the design against its requirements
Use the numbered requirements to check the final design. FR refers to the functional list; NFR refers to the non-functional list. Performance rows specify tests still required, not achieved benchmark results.
| Requirement | Design mechanism | Verification and remaining limit |
|---|---|---|
| FR1–2; NFR4: useful incident evidence | Separate signal backends, stable request/trace IDs and correct rate/histogram queries. | Trace req81, reset a counter and merge representative histograms. Verify the graph, log and trace explain the same incident without averaging instance p99 values. |
| FR2; NFR1–2: timely bounded queries | Selected indexes, bounded time ranges, appropriate resolution and isolated query resources. | Load-test the stated signal mix, thirty-second visibility and two-second dashboard p95. Include high-cardinality labels and late data instead of testing only uniform payloads. |
| FR3; NFR3–4: trustworthy alerts | Persist rule/evaluation state and gate decisions on sufficiently fresh, complete input. | Interrupt collection during a pending/firing rule. Missing evidence becomes unknown/stale, not an invented healthy zero or proof of unobserved continuity. |
| FR4; NFR1,3,5: bounded retention and safe overload | Explicit retention, bounded collector buffers, durable replay and tenant limits. | Overflow a collector, restart a writer and launch an expensive query. Report pre-acceptance losses, retain accepted backlog and verify query limits protect ingestion. |
14Rapid revision
Remember: Missing samples cannot prove zero errors. Check freshness before an alert declares recovery.
| Concern | Complete design |
|---|---|
| Metrics | Limit distinct label combinations, account for counter resets, and combine histograms with compatible buckets. |
| Logs | Give log events IDs; index selected fields and limit the scope of searches. |
| Traces | Link operation spans into traces; state the sampling policy and flag traces missing spans. |
| Collection | Batch and redact through bounded buffers so telemetry cannot exhaust the application. |
| Acceptance | Acknowledge only after the required durable save; the data may become searchable later. |
| Replay | Reuse record IDs and the backend’s documented write/retry behavior so replay does not silently duplicate data. |
| Scale | Partition using each signal’s key; keep expensive queries from exhausting collection capacity. |
| Alerting | Save alert state changes and require fresh, complete evidence; missing samples do not mean zero errors. |
| Cost | Limit new series, indexed fields, resolution and retention; disclose which detail those limits remove. |
Close with req81: a metric reveals the symptom, a trace identifies the slow dependency and a log explains the timeout. Then fail the collector and show the dashboard becoming stale. This demonstrates both the diagnostic product and the limits of its evidence.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
Why retain metrics, logs and traces separately?
Reveal a model answer
Metrics summarize aggregate behavior, logs describe events and traces link timed work across a request. Each has different identity, storage and query needs.
Interviewer follow-up
How do they connect?
Reveal the follow-up answer
Use consistent service context and request/trace IDs in the detailed evidence.
What the answer must demonstrate: Connect aggregate symptoms, individual events and request-span relationships.
Why is requestId a dangerous metric label?
Reveal a model answer
It can create a new series for every request, exhausting metadata and indexes even with modest sample bytes.
Interviewer follow-up
Where should that identity go?
Reveal the follow-up answer
Structured logs and traces, with selected indexes and retention.
What the answer must demonstrate: Explain series cardinality and place per-request identity in detailed signals.
Why compute rates before summing counters?
Reveal a model answer
Each original series can reset independently. Reset-aware rates preserve those boundaries; summing raw counters first can hide resets and create false activity.
Interviewer follow-up
Does a counter value equal new events in the last minute?
Reveal the follow-up answer
No. It is cumulative state whose change over an interval must be interpreted.
What the answer must demonstrate: Handle each counter reset before combining rates across instances.
Can instance p99 values be averaged?
Reveal a model answer
No. The summaries lose distribution shape and traffic weighting. Merge compatible histograms and compute the combined quantile.
Interviewer follow-up
What limits precision?
Reveal the follow-up answer
Histogram bucket or representation resolution and the available observations.
What the answer must demonstrate: Aggregate compatible distributions before estimating a service percentile.
Does durable acceptance mean a record is already searchable?
Reveal a model answer
Not in a buffered design. Accepted input may wait for backend ingestion, so report visibility lag separately.
Interviewer follow-up
Can that input be shed like debug data in a full collector?
Reveal the follow-up answer
No. Central acceptance establishes the stated retention obligation; tighten new admission first.
What the answer must demonstrate: Distinguish pre-acceptance shedding from accepted-data retention and visibility lag.
What should an empty recent window do to an alert?
Reveal a model answer
Apply an explicit stale or missing-data state rather than infer zero errors or automatic recovery.
Interviewer follow-up
Can an old pending timestamp prove continuity across an outage?
Reveal the follow-up answer
Only if retained observations reconstruct the interval; otherwise unseen time is not evidence.
What the answer must demonstrate: Require fresh evidence and proven duration rather than equating missing data with zero.
A writer returns at normal arrival throughput. When does its outage backlog drain?
Reveal a model answer
It does not. Completion capacity must exceed continuing arrivals, or new work must be reduced.
Interviewer follow-up
Which metric reveals recovery progress?
Reveal the follow-up answer
Oldest unprocessed or accepted-but-not-visible age, alongside the backlog size.
What the answer must demonstrate: Calculate recovery throughput minus new arrivals and track backlog age.
Does a durable queue alone prevent duplicate stored telemetry?
Reveal a model answer
No. Writers must reuse record IDs, and the backend must handle repeated writes under its documented contract.
Interviewer follow-up
Are identical timestamps sufficient identities?
Reveal the follow-up answer
No. Distinct logs can share a timestamp, and conflicts within one metric stream need a declared policy.
What the answer must demonstrate: Name stable record identity and the actual backend output/retry boundary.
Blank-page exercise · 45 minutes
Build the answer yourself
Design telemetry for a checkout timeout, then restart a counter, duplicate a log batch and lose recent ingestion while an alert is firing.
- Agree numbered functional and non-functional requirements, including operational versus audit scope, signal freshness and the acceptance/loss boundary. Then trace an application request from collection to dashboard, alert and diagnostic evidence.
- Explain metric types, cardinality, reset-aware rates and valid percentile aggregation.
- Protect applications and accepted telemetry during overload while making freshness and sampling visible.
- Trace a timeout and a concurrent request using the actual durable records.
- Review the final architecture against every numbered requirement, including measured bottlenecks, targets still needing validation and remaining failure limits.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design a metrics, logging and tracing platformReq81 timed out, then collection failed and the graph became empty. Can the alert declare checkout healthy?Recall first, then reveal
No. The graph lacks evidence; it has not measured zero errors. Mark the result stale or unknown until usable observations return.
Missing is not zero.
Return to lessonDesign a metrics, logging and tracing platformHow should several servers’ p99 latency be combined?Recall first, then reveal
Combine compatible histograms or distributions, then estimate the 99th percentile. Do not average each server’s percentile.
Combine histogram counts, then estimate the percentile.
Return to lessonDesign a metrics, logging and tracing platformWhen does a telemetry backlog shrink after recovery?Recall first, then reveal
Only when processing capacity exceeds the rate of new telemetry arriving.
Backlog needs headroom.
Return to lessonFinal revision
Summary and interview notes
Use metrics, logs and traces to investigate failures. Limit collection and query costs, and show when telemetry is missing or stale so an empty graph is not mistaken for a healthy service.
Remember these points
- Preserve each signal’s meaning.
- Limit distinct series, buffered bytes and work per query.
- Separate durable acceptance from visibility.
- Treat stale evidence explicitly in dashboards and alerts.
Interview tips
- Keep signal meaning, record identity and visibility delay distinct.
- Trace req81 from counter to log and trace, then fail collection and explain the empty graph.
Important qualifications
- Traffic and latency figures are interview assumptions, not claims about a named company's deployment.
Continue after the core interview
Explore the advanced version
The advanced lesson keeps the full detailed design. Use these sections when you want to examine the stronger requirements and failure cases.
- Custom storage publication and reclamation
A purpose-built immutable-block engine requires a deeper atomic visibility and object-lifetime protocol.
- Detailed record identity and projection metadata
High-rate bounded deduplication and storage generations need backend-specific design.
- Complete query semantics and correction paths
Pinned historical queries and late corrections extend the simpler recent-incident contract.
- Zero-loss audit requirements
Audit ingestion changes application admission and loss policies instead of inheriting debug telemetry shedding.
Technical references
- OpenTelemetry signalsOfficial definitions of metrics, logs, traces, and their roles.
- Prometheus rate and query functionsDocuments reset-aware rates and applying rate before aggregation.
- Prometheus histograms and summariesExplains quantile aggregation and distribution accuracy tradeoffs.
- OpenTelemetry Collector resiliencyOfficial description of exporter queues, retries and persistent storage; configuration defines collector crash/loss behavior.
- Grafana Tempo introductionOfficial example of a distributed tracing backend; not a claim that its implementation uses this chapter’s custom manifest protocol.
Practice marks stay in this browser.