Observability is the ability to understand a system's behavior from the signals it emits. Instrumentation produces telemetry; monitoring checks selected signals against expected conditions; diagnosis uses that evidence to investigate why behavior changed. A dashboard is one interface to those signals, not the definition of observability. OpenTelemetry primer.
For an interview-coaching service, a successful model call may still produce feedback for the wrong recording. Track the complete task and the authoritative outcome as well as individual HTTP requests.
Define the signals and identities
| Term | Meaning | Example |
|---|---|---|
| Log/event | A timestamped record of an occurrence | Recording access denied |
| Metric | A numerical measurement or aggregation | Feedback completion rate over five minutes |
| Trace | Related operations and their causal/timing relationships | Authorize → retrieve rubric → generate → validate → save |
| Span | One operation within a trace | A model attempt with start/end and status |
| Task ID | Application identity of the user's intended work | One requested feedback report |
| Attempt ID | Identity of one execution attempt | Second provider call after a safe retry |
| Operation ID | Stable identity for a consequential effect | Save the same report once across retries |
Logs, metrics and traces are complementary signals, not three boxes that automatically make a service diagnosable. A long-running task may use several linked traces across queues, pauses and retries. Do not use a trace ID as a permanent business identity or an authorization credential.
Interview scope and requirements
The service produces interactive feedback and longer background reports. Operators need to detect failures, investigate one task, reconcile costs and evaluate quality without exposing private recordings.
Functional requirements
- Connect API, retrieval, model, validation, tool and background-job operations.
- Record task outcomes separately from transport status and model completion.
- Compare behavior by release, workflow and bounded operational dimensions.
- Link sampled quality labels and billing events to the responsible task.
- Provide actionable alerts and an access-controlled investigation path.
Non-functional requirements
- Bound instrumentation overhead, exporter queues, storage and grading cost.
- Protect content, credentials and tenant boundaries in telemetry itself.
- Preserve context across services and asynchronous work.
- Make missing, sampled, delayed and dropped data visible.
- Define retention/deletion and test collection during dependency failures.
Start with task outcome counters, latency distributions, structured errors and spans for the expensive boundaries. Add payload sampling only for questions that metadata cannot answer.
What to measure
| Area | Measures | Definition to settle |
|---|---|---|
| Demand and capacity | Admitted/rejected requests, queue age, concurrency, provider quota | Accepted work versus all attempted work |
| Availability | Useful successful completions / eligible tasks | Time window; partial, deferred and rejected outcomes |
| Latency | Queue time, TTFT, completion time, inter-token gaps | Client or server start; missing first token; cancellation |
| Quality | Correctness, support, completeness, appropriate abstention | Rubric, grader, sample and label maturity |
| Actions | Authorized, committed, rejected and unknown effects | Authoritative receipt versus generated claim |
| Economics | Billed usage, attempts, tools, review and cost per successful task | Price revision, billing unit and unresolved usage |
| Telemetry health | Export failures, dropped events, lag, incomplete traces | Whether silence means healthy or unobserved |
TTFT is time to first token from a specified start point. A model-local TTFT excludes earlier queueing or retrieval; the user's TTFT includes it. For N output tokens with N > 1, (last_token_time − first_token_time)/(N − 1) estimates average inter-token latency for that response, not total completion latency. Chunked delivery can obscure individual token timing, so label the actual measurement.
Track latency distributions rather than only means. A p99 is a distribution percentile, not a maximum, and component p99s cannot generally be added into end-to-end p99. Retain enough context to distinguish slow admission from slow generation.
Trace one request without exposing everything
Record task/attempt identifiers, bounded outcomes, sizes, timings, deployment identity and model/prompt/tool/index/policy versions. Use source references for approved diagnostic content. Capture no raw prompt by default merely because an SDK supports it.
- Classify content and choose the minimum fields needed for a stated purpose.
- Redact before export where possible; test exceptions and tool payloads too.
- Enforce telemetry-store access, tenant scope, region and retention.
- Account for exports, backups, annotation sets and downstream evaluation copies.
- Audit access and monitor whether instrumentation bypasses redaction.
Hashing predictable content does not anonymize it. Propagated baggage and headers can cross service or vendor boundaries; never place secrets or unnecessary personal data there. Observable tool decisions can be traced, but a generated rationale is not guaranteed access to the model's actual internal reasoning.
Build the collection path
Read diagram source
flowchart LR
U[User task] --> A[Application and workers]
A --> C[Bounded telemetry collector]
C --> M[Metrics and SLO alerts]
C --> T[Access-controlled traces and events]
T --> Q[Durable sampled quality jobs]
Q --> E[Versioned labels and outcome reports]
E --> D[Release and incident decisions]
M --> D
T --> D
C --> H[Export lag and dropped-data signals]
Use trace-context propagation for related calls and links for asynchronous relationships where appropriate. Persist the business task ID across restarts; do not keep one application span open for days merely to represent a durable job.
Telemetry export must not create an unbounded queue during an outage. Batch with explicit limits, define which diagnostic events may be dropped, and record losses. Business-critical audit or financial records need their own durable contract; a best-effort trace exporter is not a payment ledger. Background quality grading that must survive restarts also needs durable jobs, not an untracked in-process task.
OpenTelemetry's old GenAI convention page now redirects readers to the separate GenAI semantic-conventions repository. Pin the instrumentation/convention version and check each attribute's current stability. Provider integration does not automatically supply product outcomes or make a custom field standard.
Choose metric types and avoid misleading aggregation
| Instrument | Use | Common mistake |
|---|---|---|
| Counter | Cumulative requests, errors or recorded usage | Subtracting values without handling process resets |
| Gauge | Current queue depth or in-flight work | Using the last quality score as the period's average |
| Histogram | Distribution of duration, size or scores | Buckets too coarse near an SLO boundary |
| Summary quantile | Client-computed percentile under its configuration | Averaging per-instance p95s into a service p95 |
Aggregate compatible histogram distributions first, then calculate a percentile. Prometheus currently recommends native histograms where supported; classic histograms remain useful when required by the collection path. Both need an understood accuracy/cost tradeoff. Histogram and summary guidance.
A label combination creates a distinct metric series. An illustrative 20 routes × ten model revisions × five status classes × three regions permits 3,000 combinations, before replicas and histogram bucket expansion. Adding 100,000 user IDs makes the possible cross-product enormous. Keep per-user investigation in controlled records, not ordinary metric labels. Bound retained revision labels as releases accumulate.
Sample quality without corrupting the denominator
Head sampling decides early whether to keep a trace. Tail sampling decides using completed or sufficiently collected trace information, allowing preference for slow/error cases but requiring buffering and handling late spans. Quality-review sampling is a separate decision and may occur after the task ends.
Suppose a day has 90,000 routine tasks and 10,000 difficult tasks. You review 100 of each and find failure rates of 2% and 20%. The unweighted sample average is 11%, but the traffic-weighted estimate is 0.9 × 2% + 0.1 × 20% = 3.8%, assuming representative sampling within each stratum. Report uncertainty and how each sample was selected. A targeted incident sample usually cannot estimate the total failure rate by itself.
Record eligible tasks, selected tasks, completed labels and inclusion probabilities where needed. Missing labels may be systematically harder cases. User ratings, regeneration and copy events are useful clues; they are not unbiased correctness labels. Separate input-distribution drift, system changes and grader drift during diagnosis.
Cost tracking is a reconciliation problem
Record usage per provider request/attempt and join it to the task. Keep cache-read, cache-write, uncached input, output and other billed categories non-overlapping according to that provider's contract. Include tools, retries, reasoning charges where billed, hosting and review. Missing price metadata means unknown cost, not zero.
At 100,000 tasks/day, ten retained 1 KB spans per task give about 1 GB/day of uncompressed payload before indexing, replication and overhead. A 10% representative sample reduces that portion to roughly 0.1 GB, but critical-event retention and evaluator records add cost. Reconcile estimated usage with invoices and investigate gaps. See AI FinOps.
Alerts that change behavior
Define the objective, denominator, evaluation window, owner, first action and recovery condition. Use an immediate route for severe unauthorized effects; use sustained error-budget burn for recurring availability failures and a review queue for longer cost trends.
For a 99.9% SLO, the permitted error fraction is 0.1%. A measured 1% error fraction consumes that budget at 10 times its steady rate. If sustained across a 30-day window's traffic assumptions, it would exhaust the budget in about three days. Volume changes and rolling windows affect the exact result. Multiple windows help distinguish bursts from sustained burns. SRE alerting guidance.
Do not copy universal thresholds such as “page whenever latency exceeds five seconds.” A five-second batch step and a five-second conversational pause have different consequences. Deduplicate related alerts and link a runbook that can stop a rollout, restrict actions, shed load, reconcile unknown effects or restore a compatible version.
Current tooling and the cost of each choice
| Option | Useful role | Decision to verify |
|---|---|---|
| OpenTelemetry plus metrics/trace backend | Portable application instrumentation | Exporter behavior, schema support, retention and operating cost |
| LangSmith | Traces, evaluation and development workflow | Framework-independent coverage and data controls |
| Langfuse | Traces, observations and evaluation integration | Current SDK, capture defaults and deployment responsibilities |
| Phoenix | OTel/OpenInference traces, labels and experiments | Access, scale and the actual deployment contract |
| W&B Weave | Tracing, evaluation and versioned experimentation | Workflow integration and telemetry governance |
| Helicone | LLM request visibility and usage analysis | Whether proxy placement sees every application stage |
Langfuse's current Python examples use start_as_current_observation() and observe; manual observations require explicit completion. Automatic input/output capture needs a deliberate policy. Self-hosting changes who operates the store; it does not automatically make telemetry private or secure.
Interview tip: Draw one failed user task and identify the evidence that separates retrieval, model, tool and instrumentation failures. Then explain what an on-call engineer can actually do with that evidence.
An annotated trace that answers an incident question
The following trace is invented. Times are elapsed from the root request's start, so spans can be compared without adding overlapping durations twice.
| Trace ID / span | Parent | Start–end | Safe attributes and result |
|---|---|---|---|
t42 / request |
none | 0–2,450 ms | release support-b, tenant pseudonym ta, outcome abstained_missing_evidence |
t42 / authorize |
request | 0–30 ms | policy acl-8, decision allow |
t42 / retrieve |
request | 30–230 ms | index policy-19, top_k 8, permitted_results 8 |
t42 / pack |
request | 230–250 ms | evidence tokens 6,200, required-source-present false |
t42 / generate |
request | 250–2,350 ms | immutable model revision, prompt digest, input/output/cached token counts |
t42 / validate |
request | 2,350–2,450 ms | citation validity pass, task answerability fail, safe abstention |
If the model's first token arrives at 600 ms, root TTFT is 600 ms; generation-local TTFT is 350 ms. These differ because upstream work matters. HTTP 200 can coexist with an unsuccessful user outcome. A trace reference can point to separately protected/redacted evidence for debugging without putting raw personal documents in every log.
Minimal OpenTelemetry instrumentation
This example assumes an OpenTelemetry SDK provider/exporter has already been configured. It shows the application boundary rather than a vendor-specific LLM integration. The retrieve and generate functions are application adapters.
from opentelemetry import trace
tracer = trace.get_tracer("interview.support")
def answer_request(question, release_id, retrieve, generate):
with tracer.start_as_current_span("answer_request") as root:
root.set_attribute("app.release", release_id)
with tracer.start_as_current_span("retrieve") as span:
evidence = retrieve(question)
span.set_attribute("app.evidence_count", len(evidence))
with tracer.start_as_current_span("generate"):
result = generate(question, evidence)
root.set_attribute("app.outcome", result["outcome"])
return result
Use context propagation across services and queues; without it, individual spans do not form the end-to-end trace. Record bounded metadata, apply redaction/access policy, and avoid high-cardinality IDs as metric labels. OpenTelemetry's Python instrumentation guide covers provider setup, spans, attributes, and propagation. The custom app.* fields above are application-defined, not claimed standard GenAI conventions.
Detect a change, then locate its cause
Suppose supported-answer rate falls from an observed 94% baseline to 86% after a parser release, while HTTP errors and model latency remain stable. Compare time-matched cohorts and source/language slices; check sample size, grader version, source freshness, candidate retrieval, and packed context. If table-heavy documents lost headers, roll back or repair that parser and reindex the affected snapshot. Swapping the model would not restore missing headers.
An illustrative actionable alert policy is:
signal: supported_answer_rate
window: 30m
minimum_labeled_cases: 200
condition: below_release_baseline_by_more_than_5_percentage_points
slices: [overall, table_documents]
owner: knowledge-quality-oncall
first_actions:
- inspect paired traces and recent parser/index/model/grader changes
- restrict affected corpus if answers may mislead users
- replay the regression set before expanding traffic
This is policy pseudocode, not a monitoring vendor's executable syntax or a statistical significance test. Pair it with severe-error alerts that do not wait for 200 cases. Keep all critical failure traces subject to privacy policy, plus a known-probability sample of ordinary traffic. Tail sampling improves incident visibility but biases population-rate estimates unless you account for its selection.
Interview questions with developed answers
Q1: What metrics would you track for a production LLM system?
Sample answer: I track operational health, task quality, and economics together. Operational measures include accepted and rejected demand, queue age, errors, TTFT, completion latency, and dependency limits. Quality measures depend on the workflow: supported answers, verified actions, appropriate handoff, and severe failures. Economics includes tokens, tools, retries, review, and cost per successful task. I slice by workflow and relevant version or user group, while controlling metric cardinality. Each key signal has a denominator, owner, and intended response, so the dashboard supports decisions rather than merely displaying activity.
Follow-up: Which metric proves a refund happened? An authoritative payment outcome, not the model's wording.
Q2: How do you detect quality degradation in production?
Sample answer: I combine representative sampled review, verified outcomes, targeted regression probes, and user feedback. I compare meaningful slices to a versioned baseline and inspect recent changes in prompts, models, tools, data, and graders. A shift in traffic mix can explain an average change, so I investigate examples before choosing a repair. For serious harm I contain the affected capability immediately. I then confirm the diagnosis and add regression coverage. Feedback is a useful signal, but silence or a thumbs-up is not a complete correctness label.
Follow-up: How do you separate judge drift from product drift? Regrade a stable reference set with the old and new judging configurations.
Q3: What does tracing add beyond logs?
Sample answer: Logs describe events, while a trace connects the operations that produced one task's result and shows their timing and relationships. In a RAG request I can see ingestion-version references, retrieval, reranking, generation, validation, and retries. If the needed passage was found but removed during packing, that relationship matters more than isolated success logs. I preserve task identity across asynchronous jobs and human review. Traces also help attribute cost and latency to the actual stage responsible, rather than blaming every delay on the model.
Follow-up: Must traces contain full prompts? No; metadata and controlled payload access can provide useful diagnosis with less exposure.
Q4: Your bill doubled but request count stayed flat. What do you inspect?
Sample answer: I break cost into calls per task, tokens per call, input/cache/output categories, model and service tier, tools, and retry rates. I compare traffic slices and versions. Longer histories, more agent steps, lower cache reuse, a routing change, or higher review demand can raise cost with flat top-level traffic. I link provider usage to task traces and invoices, then fix the identified cause and recheck quality. Request count is too coarse to explain an LLM bill on its own.
Follow-up: What if usage logs do not reconcile with invoices? Treat incomplete metering as a problem before trusting optimization estimates.
Q5: How do you design telemetry without leaking customer data?
Sample answer: I decide which fields are necessary for each operational purpose, avoid indiscriminate payload logging, and separate sensitive samples from routine metrics. Access controls, retention limits, redaction, and audited retrieval apply to the observability store too. I test that secrets and tenant data do not appear in traces or exports. I also account for the observability vendor and region. A debugging need is not unlimited permission to retain every prompt, document, and tool response indefinitely.
Follow-up: Can a hash still be sensitive? Yes; low-entropy or guessable content can be matched, and identifiers can link records.
60-second interview answer
I instrument the full task, not just the model call. A trace connects retrieval, model configuration, tool actions, retries, approvals, and the final outcome using a correlation ID. Metrics show quality, latency, errors, cost, and queue pressure; selected traces explain why they changed. I version the important components, minimize sensitive content, and define access and retention policies. Alerts should lead to an owner and a runbook. Observability helps diagnose behavior; it does not replace evaluation or expose a model's true internal reasoning.