System-design interview · Extended interviews
Design a metrics, logging and tracing platform
Design telemetry identities, bounded ingestion, time-series and log storage, distributed queries, sampling, retention and reliable alerts.
You will learn to
- Distinguish metrics, logs, and traces through one concrete incident.
- Calculate series cardinality, sample traffic, and log storage before indexing.
- Handle counter resets, percentile aggregation, delayed data, and monitoring failure.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Capacity estimation: throughput, latency, concurrency and storage · Database indexes: B-trees, composite keys and query access · Message queues, event logs, delivery guarantees, and backpressure
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Problem and scope
An observability platform collects metrics, logs and traces to diagnose system behavior. Metrics summarize rates and distributions; logs record individual events; traces connect spans, timed records of individual operations, across one request. Keep each signal's meaning intact while limiting the number of distinct metric series, the volume accepted, the work allowed per query and how long records are kept. Request req81 is an example diagnosis: a checkout latency metric detects a regression, a log identifies the timeout, and a trace attributes duration to inventory.
A time series is a sequence of timestamped measurements with one identity, such as requests_total{service="checkout",instance="c3",status="500"}. Each different label combination is a different series. A counter increases as events occur and may reset on restart. A gauge represents a current value, such as queue depth. A histogram records counts across value ranges so a distribution can be combined later.
I ask the interviewer whether this is a monitoring platform, a financial audit archive, or both. We choose operational telemetry with explicit loss reporting; legally required audit events use a separately specified durable path. I also ask which queries matter during an incident. The operator needs service-level error and latency graphs, a narrow search by request ID, and a trace view. The client does not need unrestricted joins over every log ever emitted.
02Functional requirements
- Query linked evidence. Authorized users query a bounded tenant/time range, open individual traces, define alert rules, and navigate from a fired alert to its supporting evidence.
- Ingest telemetry. The service acknowledges valid samples, such as checkout measurements, after durable storage; it returns a reason for each rejected sample.
- Search logs. Matching events for a request such as req81 appear with event and ingestion timestamps.
- Graph error rate. Counter resets are handled per original series.
- Graph latency p99. Compatible histogram distributions are merged before quantile estimation.
- Evaluate an alert. Pending, firing, resolved, and stale states are distinguishable.
- Apply retention. Expired data stops being queryable under a documented deletion window.
Scope and acceptance boundaries
Applications submit metric batches, structured log batches, and linked trace spans. Authorized users query a bounded tenant and time range, open an individual trace, define alert rules, and navigate from a fired alert to the evidence that produced it. An accepted batch returns a durable ingestion identity; searchable visibility follows asynchronously and has a separately measured delay.
Partial batch acceptance must be explicit. A client receives per-item errors or an all-or-nothing policy, rather than retrying every item after an ambiguous mixed response. We choose batch-level rejection for malformed envelopes and per-record outcomes for valid envelopes containing invalid records. Producers reuse record IDs on retries so accepted items are not counted twice.
Trace sampling can omit some requests; the trace UI shows that limitation. We do not infer that a missing span proves a service was never called.
03Non-functional requirements
- Ingestion latency and availability. Assume durable acknowledgment p95 below 200 ms and 99.9% ingestion availability under admitted load.
- Visibility and queries. Target recent metric visibility within 30 seconds and narrow dashboard queries p95 below two seconds.
- Alert cadence. Evaluate every 15 seconds. If required data is too old, report stale rather than a healthy resolution and use a separate platform-health route when appropriate.
- Durability. Accepted batches survive one zone failure through replicated ingestion storage. Before acceptance, unflushed collector memory may be lost under the application's configured policy.
- Retention and lateness. Use 30 days of hot metrics, seven days of indexed logs, and ten minutes of normal event-time lateness. Later data enters correction/archive processing instead of silently rewriting already evaluated recent windows.
- Security and tenant isolation. Require explicit tenants, authentication, quotas and private-data redaction. Reject or sanitize sensitive/problematic labels and fields at ingestion; dashboards cannot undo a leak or expensive index.
- Bounded resource use. Bound collector memory, query time/range and execution. Exclude unlimited queries and a universal full-text index over every payload.
Starting point and data policy
Begin with application instrumentation, a local collector, time-series database, log store and dashboard/alert process. Batch signals so each application request does not open a network connection. Support searchable structured logs and trace IDs linking related events, with lower-cost archives where appropriate.
Decide which records may be sampled and whether temporary telemetry loss is acceptable. Required security/audit records need a separate durable path. Customer emails in labels or log fields can both leak information and raise indexing cost.
Two overload contracts
What the service may drop changes once it durably accepts a record. The collector protects the application while telemetry is still waiting to be accepted; the central service has a retention obligation once it acknowledges that data.
| Before service acceptance | After service acceptance |
|---|---|
| Protect application availability with bounded buffers and reported dropping of records | Protect accepted data with durable queues and delayed visibility |
An uptime percentage does not describe either loss boundary.
04Capacity estimates
Series cardinality is the number of distinct metric-and-label combinations. Each needs metadata and index space even if it has few sample bytes, which is why the estimates below separate identity counts from ingestion bandwidth.
Use separate budgets for metric samples, log bytes, series cardinality and recovery work. Payload figures exclude envelopes and must be benchmarked against compression/index overhead.
| Estimate | Arithmetic | Consequence |
|---|---|---|
| Metric ingestion | Ten million active series / 15 s = 666,667 samples/s | A series is a distinct label set |
| Metric payload | 666,667/s × 16 B ≈ 10.67 MB/s, or 0.922 TB/day | Illustrative scalar timestamp/value payload |
| Replicated hot metrics | 30 days × three replicas ≈ 83 TB raw payload | Before indexes and compression |
| Log payload | 100,000 events/s × 500 B = 50 MB/s, or 4.32 TB/day | Seven days = 30.24 TB before replicas/indexing |
| One metric family's cardinality | 100 services × 20 regions × 10 statuses × 50 instances = one million possible series | Only if every combination exists |
| Series metadata | Ten million × 200 B = 2 GB | Before indexes, allocator overhead and replicas |
| One-hour outage | About 38.4 GB metrics + 180 GB logs | Before envelopes and replicas |
| Log recovery backlog | 100,000/s × one hour = 360 million events | Nominal capacity alone never catches up |
| Net log drain | 150,000/s recovery − 100,000/s arrivals = 50,000/s | 360 million / 50,000 = two hours |
Cardinality and indexing
Adding a million distinct user IDs makes the theoretical label space enormous. High-cardinality request IDs belong in logs/traces rather than ordinary metric labels. Limit active series and newly created series per minute separately from byte throughput: rapid creation of new identities can exhaust the series dictionary even when sample payloads are small.
Index selected structured log fields and keep large bodies in cheaper storage. Indexing every field can multiply cost.
Query shape is part of capacity
Scanning all seven days of 30.24 TB for every dashboard is unacceptable. A service/time index can reduce one request investigation to a few relevant partitions. Design that bounded query before selecting a storage product; compression and label metadata can substantially change the measured storage result.
The 16-byte sample estimate is for scalar timestamp/value pairs. Classic histogram buckets contribute multiple series; native histograms carry larger variable-size samples. Include their actual series counts and encoded sizes before using this total to size a latency-monitoring workload.
05APIs and contracts
The protocol needs to identify both a submission and the records inside it. A producer epoch identifies one run of a producer process. A batch sequence orders submissions within that run. Together they prevent a restart that reuses sequence numbers from being mistaken for old batches.
| Input or query | Example | Meaning |
|---|---|---|
| Metric batch | {"tenant":"shop","series":"requests_total","labels":{"service":"checkout","instance":"c3","status":"500"},"samples":[[t,81]]} |
Counter sample, not 81 new events |
| Structured log | {"eventId":"log81","requestId":"req81","service":"checkout","severity":"error","message":"inventory timeout"} |
One searchable event |
| Metric query | Checkout error rate over five minutes | Aggregate matching series over time |
| Log query | service=checkout AND requestId=req81 |
Find incident detail |
A batch envelope carries producerId, producerEpoch, batchSequence, and immutable record IDs. The service returns {batchId:"b81", status:"accepted", acceptedAt:..., visibility:"pending"} only after the replicated ingestion log accepts the records. Retrying an identical envelope returns the prior acceptance; conflicting content under the same identity is rejected.
Metric samples identify tenant, normalized series, event timestamp, and producer sequence where multiple legitimate observations can share a timestamp. Counter scrapes from one logical stream normally permit one value at a timestamp; conflicting repeats are rejected under a documented policy. Logs use eventId and traces use traceId/spanId. A wall-clock timestamp alone is not a universal deduplication key.
Queries specify tenant implicitly from authenticated scope, a time range, step/resolution, and maximum output size. Log pagination uses a cursor over (eventTime,eventId) bounded by a query snapshot or declared live-search semantics. Responses include dataThrough, the time boundary represented by the returned data, plus query completeness and any sampling indicator. A query budget violation returns a clear limit error with a narrower-range suggestion rather than timing out after consuming unbounded resources.
06Data model and access patterns
Accepted telemetry is transformed into queryable storage, called a projection. Writers group records into immutable blocks or chunks, then publish metadata that tells queries which blocks to read.
An input offset is a position in an ingestion-log partition; a checkpoint records the next offset a writer should consume. A manifest lists published blocks, and updating it creates a new published data generation. A separate writer ownership generation identifies which worker may publish. These records connect durable input, restart progress and searchable output.
| Store | Key and content | Authority |
|---|---|---|
| Ingestion log | partition and offset; accepted records | Durable source for the replay window |
| Series dictionary | tenant + canonical metric and labels | Maps identity to a series ID |
| Time chunks | series ID and interval | Published immutable sample blocks |
| Log blocks and selected indexes | tenant, time bucket, event ID | Searchable event representation |
| Trace spans | trace ID, span ID, parent ID | Request relationship representation |
| Chunk manifest and checkpoint | partition, generation, published objects, next offset | Which stored output is visible |
| Alert state | rule ID, evaluation boundary, pendingSince, status | Durable evaluation and notification state |
Event time records when checkout observed a timeout; ingestion time records when the platform accepted it. Keeping both enables lag diagnosis. A sample may be old without the ingestion system being slow if the originating machine buffered it for hours.
The ingestion log is retained long enough to repair normal writer failures. Immutable data blocks remain in hot storage and later object storage according to retention. Queries can read an uploaded block only after its reference is published in the metadata manifest. Derived indexes must correspond to that same generation or declare their lag, avoiding a response that claims completeness while silently omitting newly published records.
Shard metrics by tenant and series identity, and logs by tenant and time plus a distribution key. A huge tenant receives multiple partitions; one tenant ID must not force all of its traffic onto one writer.
Writers must save record-deduplication state durably; an in-memory set disappears when the worker crashes. Route every repeated record identity to the same owner, retain its fingerprint and outcome for a declared retry window, and advance that deduplication state with chunk publication/checkpointing. Otherwise a repeated record at a later log offset could be counted again in another chunk, even if replay of each offset range is safe. Size this metadata separately; for example, 100,000 distinct log IDs/s retained for 24 hours is 8.64 billion identities. A workload with this volume may use a shorter bounded retry window or a proven producer-sequence protocol, but must state the resulting client contract.
07Basic working design
The first working system instruments checkout, batches telemetry in a local collector, stores metrics in a time-series engine and logs in a structured event store, and runs a query/alert service. On req81, the application increments its request counter, observes the duration histogram, and emits a timeout log carrying the trace ID. The collector submits a bounded batch and releases its local copy after the service acknowledges durable acceptance.
At low volume, the storage engine can be the durable acceptance boundary without a separate broker. The query engine selects checkout's series and reads five minutes of samples. An alert calculates the error ratio and requires it to remain above the configured threshold for a stated duration before firing. That duration filters transient noise but adds detection delay; it should be chosen with the on-call response objective.
The baseline redacts private fields before they enter persistent storage. It also bounds collector memory and reports rejected or dropped telemetry. These are part of a functional product, not optional features deferred until scale. Otherwise a monitoring outage can cause the checkout outage it is supposed to explain.
The operator opens the alert's fixed evaluation range, follows req81's trace, and narrows logs to the inventory timeout. The system has now completed a useful end-to-end incident workflow.
The collector batches signals; the alert links aggregate symptoms to request detail.
Read each connection in order
- syncSamples + req81 log / traceInstrumented applications → Bounded local collectors
- syncBatch ingestionBounded local collectors → Metric and structured log stores
- syncBounded metric / log queryDashboard / alert process → Metric and structured log stores
- syncAlert and evidence linksDashboard / alert process → On-call user
08Find the baseline flaws
The baseline receives about 60.67 MB/s of raw metric and log payload before envelopes. A shared database node serving a broad historical search can consume its disk bandwidth and delay current ingestion. During an incident, precisely when the operator opens more dashboards, monitoring freshness worsens. Scaling the query process alone does not isolate shared disks.
Correctness can fail even when every write succeeds. Averaging two instance p99 values does not produce the service p99. Summing counters before handling resets makes a restarted instance look like negative traffic. An empty input window can falsely resolve an alert if the evaluator treats missing observations as zeros. The stored samples may be intact; the calculations interpret them incorrectly.
A crash creates another counterexample. Writer A uploads a new chunk containing offsets 118–130, then crashes before recording its checkpoint. A replacement replays those offsets. If both chunks become visible without an atomic publication rule, log counts and histogram buckets double. A queue's delivery guarantee cannot by itself make the derived storage exactly-once.
Finally, an attacker or accidental instrumentation change adds requestId as a metric label. Cardinality can grow with every request despite a modest byte rate. Byte quotas alone do not protect series-index memory.
09Improve the design, step by step
First, insert a replicated ingestion log. The trigger is storage downtime or bursts exceeding writer capacity. Gateways validate and append accepted records; independent writers consume them. The service can accept records while searchable storage is delayed, and writers can replay the log after failure. The cost is additional storage, visibility delay, and a finite backlog budget. Direct-to-store ingestion remains preferable at small scale when the engine already provides adequate buffering and recovery.
Second, partition identities and publish immutable chunks. The trigger is the aggregate sample rate and replay cost. Series-key partitions preserve useful local order while multiple writers build chunks. Writers publish a chunk manifest and input checkpoint atomically under a current ownership generation. This improves parallel throughput and restart safety. It adds metadata coordination and orphan cleanup; a stale writer may upload bytes but cannot publish them. Mutable per-sample rows remain attractive for lower throughput or frequent corrections.
Third, separate query resources from ingestion. The trigger is a historical search delaying recent telemetry. Query workers read published chunks and selected indexes from dedicated capacity, with per-tenant scanned-byte and concurrency budgets. Ingestion latency becomes less sensitive to investigative queries. The cost is duplicated caches and possible query throttling. A single engine remains simpler when measured query demand is small and predictable.
Fourth, tier retention and bound cardinality. The trigger is multi-terabyte daily logs and series churn. Recent chunks stay on fast storage; older immutable blocks move to object storage. Aggregates retain the counts, sums and histogram information needed by supported queries. Admission limits active series and new series creation. This lowers hot storage cost, but older queries become slower and downsampling loses temporal detail. Keeping raw high-resolution data is appropriate for a short critical incident window or a separately funded audit requirement.
These changes do not justify indexing every field. The query contract still chooses low-cardinality labels for metrics and selected structured fields for logs. Adding another service does not reduce series cardinality when the data model gives every request its own time series.
10Detailed architecture
Collection and durable ingestion
Application libraries emit three signal types into collectors with bounded local buffers. Authenticated ingestion gateways validate schemas, assign tenant scope, enforce byte and cardinality budgets, and append to partitioned replicated logs. Collector retry is tied to durable acceptance, not dashboard visibility.
Storage publication and queries
Each storage writer holds a log partition’s current ownership generation and builds immutable metric or log blocks. A metadata authority publishes block manifests with their checkpoints; its replicas protect that commit boundary. The series dictionary and selected log indexes support bounded lookup. Object storage holds durable blocks, while recent hot caches accelerate dashboards without becoming the authority for accepted data.
Query workers authorize a tenant and select one coherent manifest generation, then read the necessary chunks and indexes. Rule evaluators issue bounded recent queries, check completeness and freshness, and persist transitions before notifying. A notification service uses stable transition identities so evaluator retries do not create repeated identical pages.
Freshness and independent monitoring
Collector-to-log acceptance is synchronous. Projection into storage, indexing, compaction, and archival are asynchronous. Alerting necessarily observes a delayed view, so its response includes the data boundary it evaluated. An independent external probe watches the platform's heartbeat and freshness; relying only on this same ingestion path would make its complete failure invisible.
Trace identity and sampling
Trace spans use a tenant-scoped traceId and spanId, and a trace lookup gathers that trace’s spans from trace-oriented partitions rather than scanning every metric series. This exercise ingests completed immutable spans; repeated identical IDs are deduplicated, while a conflicting span is rejected or quarantined under an explicit policy. Missing child spans leave an incomplete trace. Head sampling makes a consistent decision near request start and propagates it; tail sampling buffers spans and decides later, for example to keep error traces, but costs memory and cannot guarantee a complete trace after its wait deadline. Late spans must not silently turn a sampled partial trace into claimed complete evidence.
Implementation option and limits
A coherent starting stack uses OpenTelemetry SDKs and collectors, Prometheus-compatible metric storage/query semantics and a trace backend such as Tempo; choose a structured log backend for the required indexed fields. Configure bounded queues and persistent exporter storage where needed. These products do not automatically implement this chapter’s custom atomic-manifest protocol: either use a backend’s documented durability/publication behavior or build that boundary explicitly, then test it under replay.
Queries read only blocks listed in the committed manifest. Separate query capacity prevents historical searches from consuming the resources reserved for ingestion writers.
Read each connection in order
- sync1. Metric / log / trace batchesInstrumented applications → Bounded collectors
- sync2. Retry stable identitiesBounded collectors → Ingestion auth / quota gateways
- sync3. Durable accepted appendIngestion auth / quota gateways → Replicated ingestion partitions
- async4. Consume offset rangesReplicated ingestion partitions → Generation-owned storage writers
- syncUpload immutable blocksGeneration-owned storage writers → Immutable blocks and index fragments
- sync5. Publish objects + checkpointGeneration-owned storage writers → Manifest + checkpoint authority
- syncChoose published generationTenant-bounded query workers → Manifest + checkpoint authority
- syncRead selected chunks / indexesTenant-bounded query workers → Immutable blocks and index fragments
- syncCache versioned recent dataTenant-bounded query workers → Recent chunk cache
- sync6. Authorized bounded queriesOn-call dashboards → Tenant-bounded query workers
- syncEvaluate fresh complete windowFreshness-aware rule evaluators → Tenant-bounded query workers
- syncPersist alert transitionFreshness-aware rule evaluators → Alert state and transition outbox
- asyncStable notification intentAlert state and transition outbox → Notification delivery
- syncAlert + investigation linksNotification delivery → On-call dashboards
- syncEnd-to-end heartbeat writeIndependent external probe → Ingestion auth / quota gateways
- syncCheck heartbeat freshnessIndependent external probe → Tenant-bounded query workers
11Write path and acknowledgement
Acceptance, buffering and durable retention have explicitly different guarantees for ordinary telemetry and required audit events. Backpressure slows or rejects new submissions; any records the system drops must be counted and reported separately.
Checkout instance
c3handlesreq81, records a duration observation, increments its counter, and emits a structured timeout log with the same trace ID as the inventory call.The local collector batches data and records it in a bounded local write-ahead buffer if that durability is required. After the ingestion gateway authenticates tenant
shop, validates labels, and durably accepts a batch into the ingestion log, it acknowledges that boundary.Consumers route samples by series ID to storage writers. Writers append to time chunks and update label indexes. Duplicate batch retries are handled by a defined sample/event identity policy; two legitimate events sharing a timestamp must not accidentally collapse.
A dashboard query first selects checkout series, then reads five minutes of chunks. The operator computes each instance counter's rate before summing across instances. This handles a reset on
c3without mistaking it for a drop in total system traffic.An alert rule observes a sustained error ratio above its threshold, enters pending state, then fires after its configured duration. The notification includes a dashboard range and trace/log links so the operator can inspect
req81, rather than a graph with no investigative path.Writer W reads a bounded offset range, applies the record-identity policy, and uploads a content-addressed block plus its index fragment. It publishes their manifest and advances its input checkpoint in one metadata transaction. Only published generations enter queries.
If acceptance reaches the collector but the searchable projection is delayed, req81 remains durable in the ingestion log. The API exposes visibility lag instead of claiming the record has vanished. If the collector loses its reply, retrying the same producer identity returns or recreates the same logical records within the supported deduplication window.
Compaction merges published blocks into a new generation and swaps the manifest atomically. Queries pinned to the old generation can finish before garbage collection removes its objects. Compaction does not alter record identities or silently add the same histogram observations twice.
A collector's bounded local disk buffer and the central log have different failure scopes. The former can survive an agent restart if configured; the latter protects data only after durable service acceptance.
12Read and delivery path
Queries use bounded time ranges and label/search scopes, and disclose missing partitions, sampling and freshness.
For long retention, keep summaries that support the required calculations: count and sum for means, appropriate bucket counts for histograms, and reset-aware information for counters. A five-minute average cannot answer every later subsecond incident question. Show the resolution and missing-data status in the UI.
The operator's request proceeds through a specific path. The query service authenticates shop, selects a fixed manifest version, selects checkout series using the label index, and loads chunks for the last five minutes. It applies each original series' counter-reset logic before combining rates. For latency it combines compatible bucket counts with their observation counts, then computes an approximate quantile with a stated resolution.
For req81 logs, the query restricts tenant, service, and event-time range before looking up the event or trace identifier. It returns the selected records with ingestion timestamps, making delayed arrivals visible. Pagination remains attached to the fixed boundary so compaction or new ingestion cannot duplicate or skip records across pages.
The response includes dataThrough and completeness. A cached dashboard result is keyed by query, tenant, resolution, and relevant data generation. Late-arriving corrections invalidate or version the affected range. Without that boundary, caching could make a repaired ingestion window continue to look empty.
13Correctness deep dive
Suppose partition p3 has committed checkpoint 118, meaning offsets below 118 are already published. Writer A at generation 7 uploads block C42 for offsets 118–130, then pauses. A new owner B acquires generation 8 and replays that same range. Uploaded objects are not yet authoritative.
publish(partition=p3, generation=8, expectedNext=118,
objects=[C43,index43], next=131):
begin metadata transaction; lock partition p3
require currentGeneration == 8 and checkpoint == 118
verify uploaded objects are complete and checksummed
append manifest entry for range [118,131)
set checkpoint = 131
commit
B's transaction installs C43 and checkpoint 131 together. When A resumes, its generation-7 publication fails even if its bytes are perfectly valid. C42 is an orphan eligible for delayed cleanup. If A had published before ownership changed, B would observe checkpoint 131 and not publish the old range again. The metadata transaction decides which publication can commit first.
A query reads a manifest snapshot rather than listing every object in the bucket. It therefore sees either the old published boundary or the newly committed range, never both C42 and C43. The index fragment is part of the same published generation, so a response cannot claim complete search visibility while using an index that omits the new block.
Object reclamation uses that same metadata authority. Each uploaded block has a staged registry row tied to its writer generation. Publication atomically transfers it to manifest references and rejects any row already marked deleting. Cleanup cannot mark an object deleting while its current staging grant, a retained manifest reference, or a live query snapshot pin—a record protecting objects still being read—protects it. Queries acquire their manifest pin transactionally before reading objects; compaction retains old references until those pins release. A crashed query pin expires only under a defined reader contract that prevents further reads without renewal. If cleanup marks an abandoned object deleting first, late publication fails; if publication or a query pin wins first, cleanup cannot select it. Waiting longer before cleanup can reduce races, but the atomic metadata checks are what prevent deletion during publication or a protected read.
The current generation publishes the offset range once; the old object never enters a query manifest.
Read each connection in order
- syncUpload C42 for offsets 118–130Writer A / generation 7 → Object storage
- syncPause before publicationWriter A / generation 7 → Writer A / generation 7
- syncAcquire generation 8; read next 118Writer B / generation 8 → Manifest authority
- syncUpload C43 for the replayed rangeWriter B / generation 8 → Object storage
- syncAtomically publish C43 + next 131Writer B / generation 8 → Manifest authority
- returnCommittedManifest authority → Writer B / generation 8
- syncPublish C42 using generation 7Writer A / generation 7 → Manifest authority
- blockedReject obsolete generationManifest authority → Writer A / generation 7
- syncRead published manifestQuery worker → Manifest authority
- returnRange references C43 onlyManifest authority → Query worker
14Failure and recovery
Lost collection, delayed storage publication and incomplete queries can all leave a dashboard without recent values. They require different recovery actions, so each failure must identify where progress stopped and what evidence remains available.
| Failure or race | Required response and boundary |
|---|---|
| Collector buffer fills | During an ingestion outage, collectors buffer only to a configured byte/time limit, retry with jitter, and report dropped data. When the buffer fills, choose documented shedding rules, such as dropping debug logs before higher-priority operational records. Required security/audit events need their own durable admission and retention contract; do not silently apply debug-log sampling to them. |
| Out-of-order or incomplete data | Partition the ingestion log by tenant and series key to preserve useful local order while distributing work. Storage accepts out-of-order samples only within a declared lateness window or routes them through a correction path. Logically duplicate samples with conflicting values require a deterministic reject/replace policy. An alert evaluating incomplete data must distinguish “no observations” from “healthy zero errors.” Track watermark or ingestion lag and expose stale evaluation state. |
| Expensive query overload | Separate ingestion capacity from expensive historical queries. Apply tenant query budgets, time-range limits, result limits, and cancellation. Use hot local chunks for recent dashboards and object storage for older data; caching common queries reduces repeated scans but must account for late-arriving data. Monitor the observability platform through an independent small heartbeat and external probe so a failed main pipeline cannot hide its own outage. |
| Ingestion quorum or writer loss | If the central ingestion log loses its write quorum, gateways stop acknowledging new records. Collectors buffer to their configured byte and time limits, then apply declared shedding. Previously acknowledged offsets remain recoverable under the stated replica-failure assumption. If storage writers fail while the log stays healthy, acceptance can continue only while reserved backlog space remains; gateways limit new acceptance before log retention could overwrite accepted data that writers have not yet published. |
| Alert evaluator restart | After a restart, an alert evaluator restores its pendingSince and last evaluated boundary. It does not restart a ten-minute pending timer on every process crash or send a fresh page for the same transition identity. If an incomplete window cannot support the rule, it remains stale rather than manufacturing a firing or resolved conclusion. |
| Missing continuity during an outage | Persisting pendingSince does not prove the threshold held during an unobserved outage. After a restart, the evaluator either reconstructs a complete qualifying interval from retained samples or resets the continuous-duration timer when continuity is unknown. A stale period cannot count as ten minutes of demonstrated failure or as a healthy resolution. A rule-generation change also starts a deliberately new evaluation identity rather than inheriting the old rule’s pending clock. |
15Operations, security, and cost
Restrict ingestion credentials to a tenant and signal class. Redact tokens and private payload fields before persistence, and audit changes to retention and alert routes. A label allowlist prevents accidental request IDs from becoming metric dimensions; rejection metrics and sampled diagnostics explain which instrument caused the problem.
The main service indicators are acceptance latency, accepted-to-queryable delay, series churn, age of the oldest accepted record still waiting for publication, collector drop count, query scanned bytes, and alert data freshness. Availability of the HTTP endpoint alone can look excellent while every alert is evaluating stale data. The external heartbeat therefore exercises an end-to-end write and read with a known timestamp.
Roll out a new label canonicalization rule with a dual-read or explicit version migration; otherwise the same metric can split into two identities. Replay representative traffic through a new writer and compare counts, sums, and histogram buckets. Recovery tests crash a writer after object upload but before publication, pause an old owner through a generation change, and replay a collector batch after a lost response.
The seven-day raw log footprint is 30.24 TB before replicas and indexes. Indexing only 20% of a 500-byte envelope would still create about 6.05 TB of indexed field payload over that window before index overhead. This does not prove a storage bill, but it makes selective indexing and retention the first cost questions. Evaluate query speed improvements against the incident evidence they discard; a faster dashboard is less useful if operators can no longer diagnose the failure.
16Decision ledger and limitations
The platform is optimized for bounded incident queries. Its main cost controls—selected labels, selected indexes and limited retention—also determine which questions an operator can answer later.
| Decision | Benefit | Cost |
|---|---|---|
| Low-cardinality metrics | Cheap fast aggregates | Per-request detail lives elsewhere |
| Structured logs and trace IDs | Targeted investigation | Instrumentation and retention work |
| Durable ingestion log | Replay after storage failure | Buffering cost and operational complexity |
| Separate query budgets | Protect ingestion during incidents | Some broad queries are delayed/rejected |
Durable acceptance costs log capacity and creates a visibility lag. Immutable chunk publication simplifies crash recovery but makes very late corrections more expensive. Selected log indexes make ordinary incident queries fast while broad searches may require queued scans. Downsampling saves storage at the permanent cost of temporal detail; the UI must show the resolution.
We do not claim every operational event survives a disconnected collector whose configured buffer overflows. We do claim that loss is counted and surfaced, and that already accepted central records follow the stated durability policy. Security audit events need their own stronger admission path if their loss is unacceptable.
The next redesign trigger is either hot-tenant skew that defeats current partitioning, or query demand that scans more bytes than the isolated query fleet can afford. The response is finer partitioning or explicit asynchronous search, not allowing unlimited queries to compete with ingestion.
17Interview closing
“The platform supports incident diagnosis: metrics identify aggregate symptoms, traces connect work across services, and structured logs explain individual requests. I keep those signal models distinct and expose their freshness. Collectors batch into bounded buffers; authenticated gateways durably accept records into a partitioned log; writers publish immutable blocks and checkpoints atomically. Queries and alerts read one published data version, using resources reserved separately from ingestion.
“The hard failure case is a writer that uploads a block and dies before checkpointing. A generation and expected checkpoint are checked in the manifest transaction, so replay exposes one range once while stale uploads remain unreferenced. Producer record identities handle duplicate submission separately. On the query side I calculate counter rates before aggregation and merge histogram counts before estimating service percentiles.
“The main costs are log bytes, active-series cardinality, indexes, and retained resolution. My next measurement is accepted-to-visible lag during a large incident search. The recovery test combines a storage outage, a writer takeover, and a stale alert window to prove we show delayed or missing data instead of a falsely healthy dashboard.”
If the interviewer changes the product into a zero-loss audit archive, I would revise application admission, offline buffering, retention, and recovery guarantees explicitly. Operational debug-log shedding would no longer be an acceptable inherited policy.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
What is the difference between a metric, a log, and a trace?
Reveal a model answer
A metric summarizes a measured quantity over time, a log records an event, and a trace relates timed steps within a request. For slow checkout I use a latency metric to detect it, a trace to locate the slow dependency, and a log for the specific error.
Interviewer follow-up
Why not use logs for every graph?
Reveal the follow-up answer
It is possible, but repeatedly scanning detailed events is more expensive than maintaining bounded aggregate series for common operational questions.
What the answer must demonstrate: Tie each signal to a concrete question.
Why is userId a dangerous metric label?
Reveal a model answer
Each distinct label combination becomes a separate time series. A user ID can multiply series count and index memory even when individual samples are tiny. I would keep it in access-controlled logs or traces instead.
Interviewer follow-up
Is low ingestion byte rate enough to make it safe?
Reveal the follow-up answer
No. Series churn and index cardinality can exhaust resources independently of payload bandwidth.
What the answer must demonstrate: Count identities as well as bytes.
How do you calculate total request rate across restarting instances?
Reveal a model answer
I calculate a reset-aware rate for each instance’s counter, then sum the rates. If I sum counters before taking the rate, a reset can be hidden by other instances or distort the result.
Interviewer follow-up
What if an instance has no recent samples?
Reveal the follow-up answer
I expose missing or stale data according to policy; I do not automatically assume its rate is zero.
What the answer must demonstrate: Preserve the original counter identity until reset handling.
Can you average ten instance p99 values?
Reveal a model answer
No. Those values do not retain each distribution or its traffic weight. I aggregate compatible histogram buckets or another mergeable distribution representation, then estimate the service quantile.
Interviewer follow-up
What accuracy limitation remains?
Reveal the follow-up answer
Histogram resolution and bucket placement bound quantile precision. I choose them for the latency range and decision being made.
What the answer must demonstrate: A percentile is not an additive measurement.
The ingestion service is down for an hour. What happens to application logs?
Reveal a model answer
Collectors use bounded buffers and retries. When the configured limit is reached, they follow explicit priority/drop policy and report loss; required audit records need a separately designed durable path.
Interviewer follow-up
Why not block every application request until logs upload?
Reveal the follow-up answer
That can turn a monitoring outage into a product outage. Only a requirement explicitly demanding that coupling justifies it.
What the answer must demonstrate: State the loss and backpressure contract.
The error-rate chart is flat at zero during a storage outage. Is the service healthy?
Reveal a model answer
We do not know. No data and zero errors are different states. The alert evaluator checks ingestion freshness and marks its result stale or unknown; an independent probe can detect the monitoring outage.
Interviewer follow-up
What would you alert on for the platform itself?
Reveal the follow-up answer
I monitor ingestion lag, dropped samples, series churn, query failures and evaluation freshness through an independent path. After an evaluator outage I reconstruct the complete threshold interval or reset its pending timer; elapsed downtime alone is not proof the rule continuously held.
What the answer must demonstrate: Missing telemetry must not become false reassurance.
A writer uploaded a chunk and crashed before advancing its offset. Why will replay not double the count?
Reveal a model answer
Queries use published manifests. The replacement writer atomically installs the chunk reference and checkpoint under the current generation and expected offset. The old uncommitted object is invisible, and the old writer cannot later publish after its generation is replaced.
Interviewer follow-up
Does this remove duplicate logical events that appear at two offsets?
Reveal the follow-up answer
No. That requires stable producer record identities and a defined deduplication window. Offset publication handles replay of processing, not arbitrary duplication at ingestion.
What the answer must demonstrate: Distinguish log replay idempotency from event identity.
Storage is down for one hour at 100,000 log events/s. It returns with capacity for 150,000/s. When is the backlog gone?
Reveal a model answer
The outage accumulated 360 million events. New arrivals still consume 100,000/s, leaving 50,000/s for recovery, so draining takes 7,200 seconds or two hours, assuming those rates remain stable. I also check log retention and disk reserve cover that interval.
Interviewer follow-up
What if capacity returns at exactly 100,000/s?
Reveal the follow-up answer
The backlog never shrinks under the same arrival rate. I need temporary excess capacity, reduced admitted traffic, or an explicitly changed freshness target; declaring the service recovered because writers are running is misleading.
What the answer must demonstrate: Use net drain rate, not gross processing rate.
Blank-page exercise · 45 minutes
Build the answer yourself
Design an observability platform that helps the operator investigate req81, then overload ingestion while one checkout instance restarts.
- Separate metrics, logs, traces, and their retention.
- Calculate samples, series cardinality, and log bytes.
- Trace collection through one alert and investigation.
- Handle counter resets and percentile aggregation correctly.
- Explain no-data, buffer limits, and independent monitoring.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design a metrics, logging and tracing platformWhere should request IDs live?Recall first, then reveal
In logs and traces; making every request ID a metric label creates a new series per request.
Metrics summarize; traces identify.
Return to lessonDesign a metrics, logging and tracing platformHow do you aggregate resettable counters?Recall first, then reveal
Compute a reset-aware rate for each original series, then sum rates.
Rate first, sum second.
Return to lessonDesign a metrics, logging and tracing platformCan you average machine p99 values?Recall first, then reveal
No. Aggregate compatible distribution data and calculate a quantile from the combined distribution.
Merge distributions, not percentiles.
Return to lessonFinal revision
Summary and interview notes
An observability platform preserves the meanings of metrics, logs and traces while separating durable ingestion from query visibility. Stable record identities and atomic projection publication protect counts; freshness-aware queries and alerts prevent missing data from appearing healthy.
Remember these points
- Count active series and churn separately from byte throughput; request IDs usually belong in logs and traces.
- Compute reset-aware counter rates per original series before summing.
- Merge compatible distributions before estimating aggregate percentiles; machine p99 values cannot be averaged.
- Commit block references, deduplication progress and checkpoints together. Protect unfinished uploads and active reads from cleanup with staging grants and reader pins.
- An alert’s pending duration requires complete supporting data, not simply an old persisted timestamp.
Interview tips
- Calculate sample rate, retained bytes, series metadata and net recovery capacity independently.
- Trace one request from instrumentation through durable acceptance, visible blocks and an alert.
- Explain both how replay avoids publishing an offset twice and how repeated event IDs at different offsets avoid being counted twice.
Important qualifications
- Collector loss before service acceptance follows the configured buffer policy; required audit events need a distinct contract.
- Sampling and downsampling discard information; response completeness and resolution must remain visible.
- Product backends supply their own protocols; the illustrative publication algorithm is not a universal vendor guarantee.
Technical references
- OpenTelemetry signalsOfficial definitions of metrics, logs, traces, and their roles.
- Prometheus rate and query functionsDocuments reset-aware rates and applying rate before aggregation.
- Prometheus histograms and summariesExplains quantile aggregation and distribution accuracy tradeoffs.
- OpenTelemetry Collector resiliencyOfficial description of exporter queues, retries and persistent storage; configuration defines collector crash/loss behavior.
- Grafana Tempo introductionOfficial example of a distributed tracing backend; not a claim that its implementation uses this chapter’s custom manifest protocol.
Practice marks stay in this browser.