Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Case Study: Clinical Voice Documentation

By Anup Rai13 min readReviewed September 2026

This is a hypothetical interview scenario. Workload, latency, staffing and costs are planning assumptions, not measured results. Model and interoperability facts have primary-source links. Clinical review and hospital-approved workflows remain part of the design.

Interview focus: produce faithful draft notes from speech, preserve patient and encounter identity, and file only an authorized version through the EHR's actual interface.

60-second interview answer

I would begin with clinician-initiated dictation for a confirmed patient encounter, a versioned transcript and an editable draft note. Streaming transcription gives immediate feedback, while stable segments feed clinical extraction and note generation. Every proposed fact retains its speaker, time, negation and source. Clinicians resolve uncertain content and approve the exact note before filing. The EHR adapter verifies permissions and reconciles uncertain writes. I would expand to ambient multi-speaker recording only after measuring attribution and omission errors. Success is accurate documentation with less correction time, including privacy, review and integration costs.

Remember: Confirm encounter → Capture → Preserve meaning → Review → File exact note.

Interview problem and scope

A hospital network wants nurses to document encounters by voice instead of retyping notes. Design a service for 10,000 encounters per month that supports noisy rooms, medical terminology and integration with Epic and Oracle Health.

First distinguish two products. Dictation records a clinician intentionally describing a visit. Ambient documentation records a conversation and must handle patients, clinicians, family members, interpreters and overlapping speech. Start with dictation; make ambient capture an explicit later requirement. Neither product should independently diagnose, prescribe or place orders.

Functional requirements

  1. Start a recording only after the authenticated clinician confirms the patient, encounter and required recording permissions.
  2. Stream provisional text, preserve finalized segment revisions and visibly report capture gaps.
  3. Extract stated clinical facts with supporting transcript/audio locations and unresolved ambiguity.
  4. Draft the hospital's approved note template; support edits, comparison with source evidence and clinician approval.
  5. File an approved version through the configured EHR integration and show pending, confirmed or failed delivery.
  6. Support corrections and amendments without silently rewriting a previously signed clinical record.

Nonfunctional requirements

  1. Meaning fidelity: evaluate medication, allergy, negation, temporal and speaker errors independently of overall transcription accuracy.
  2. Responsiveness: target first provisional text under 500 ms p95 from the agreed available-speech point; measure finalization and note generation separately.
  3. Identity safety: bind all artifacts and writes to the server-verified patient and encounter; a spoken name cannot switch charts.
  4. Privacy: approve every processor of protected health information, including audio storage, ASR, note models, support tooling and logs.
  5. Resilience: keep the existing manual documentation workflow available during model or EHR outages; never fabricate missing audio.
  6. Auditability: record source revisions, model/configuration versions, clinician edits, approvals and external write receipts.

Definitions to establish before the diagram

Term Standard meaning Boundary in this design
ASR Automatic speech recognition converts speech to text It can mishear words, numbers and negation
VAD Voice activity detection identifies speech intervals It does not identify a speaker or determine clinical importance
Diarization Separates audio by speaker identity labels “Speaker 2” does not establish patient or clinician role
NER Named entity recognition locates entities such as a drug or symptom Relations, certainty, temporality and negation require additional context
EHR Electronic health record The destination has its own permissions, templates and signing workflow
FHIR HL7's standard for exchanging healthcare information Standard-shaped JSON does not guarantee an enabled write API

SOAP is one possible template: Subjective contains reported symptoms, Objective contains observations and measurements, Assessment contains the clinician's documented assessment, and Plan contains the clinician's stated next steps. Use the institution's nursing template when SOAP is inappropriate. A missing assessment is not an invitation for the model to invent one.

Capacity and latency estimates

Assume 15 recorded minutes per encounter and 20 eight-hour working days per month.

Quantity Calculation Implication
Recorded minutes 10,000 × 15 = 150,000/month Bill transcription by actual processed audio, including retries
Stream hours 150,000 / 60 = 2,500/month Mean concurrency over 160 working hours is 15.6 streams
Peak concurrency Five times mean ≈ 79 streams Benchmark at peak plus reconnect and redundancy headroom
Raw mono audio 16,000 samples/s × 2 bytes × 900 s × 10,000 288 GB/month before replicas and derived artifacts
Compressed audio example 64 kilobits/s × 150,000 minutes 72 GB/month in decimal units; codec support and quality must be validated

A machine that transcribes one completed recording at ten times real time does not necessarily support ten concurrent streams at a tight tail-latency target. Measure streaming batching, memory, scheduling and failover behavior with representative clinical audio.

Clock Illustrative target Start and end
First provisional text Under 500 ms p95 Enough speech available to the client → first displayed partial
Stable utterance Under 1 second p95 Actual end of utterance → final segment displayed
Draft update Under 3 seconds p95 Stable transcript revision → corresponding draft ready
Approved note Measured, no model-only promise Encounter documentation start → clinician approval
EHR confirmation Contract-specific Approved submission → verified external record

A first-partial planning budget might allocate 50 ms to framing, 60 ms to transport, 80 ms to queueing, 200 ms to recognition and 50 ms to rendering: 440 ms total. These component allowances are not measured percentiles; adding individual p95 values does not produce an end-to-end p95. Endpointing—the decision that a speaker has finished—adds a separate delay.

Baseline, then detailed design

The baseline is push-to-talk dictation, an approved transcription engine and a manually edited note. Bind the recording to the selected encounter, retain revision history and require normal clinician approval. This yields value without introducing multi-speaker role inference or autonomous clinical reasoning.

Add structured extraction, evidence-linked drafting and ambient recording only when they improve measured documentation effort and quality.

Architecture / visual model
flowchart TB UI[Clinician confirms patient encounter and recording] --> CAP[Capture with pause indicator and sequence numbers] CAP --> STREAM[Authenticated audio stream] STREAM --> ASR[Approved streaming ASR] STREAM --> SPEAK[Optional audio diarization] ASR --> ALIGN[Versioned words and speaker alignment] SPEAK --> ALIGN ALIGN --> FACT[Clinical facts with evidence and uncertainty] ALIGN --> DRAFT[Template-aware note draft] FACT --> DRAFT DRAFT --> CHECK[Source fidelity and missing-content checks] CHECK --> REVIEW[Clinician evidence and edit screen] REVIEW --> APPROVE{Approve exact revision?} APPROVE -->|Edit| EDIT[New draft revision] EDIT --> REVIEW APPROVE -->|Yes| OUTBOX[(Approved submission intent)] OUTBOX --> ADAPT[EHR adapter and reconciliation] ADAPT --> EHR[(Epic or Oracle Health record)] STORE[(Restricted source and audit storage)] --- ALIGN STORE --- OUTBOX
Read diagram source
flowchart TB
    UI[Clinician confirms patient encounter and recording] --> CAP[Capture with pause indicator and sequence numbers]
    CAP --> STREAM[Authenticated audio stream]
    STREAM --> ASR[Approved streaming ASR]
    STREAM --> SPEAK[Optional audio diarization]
    ASR --> ALIGN[Versioned words and speaker alignment]
    SPEAK --> ALIGN
    ALIGN --> FACT[Clinical facts with evidence and uncertainty]
    ALIGN --> DRAFT[Template-aware note draft]
    FACT --> DRAFT
    DRAFT --> CHECK[Source fidelity and missing-content checks]
    CHECK --> REVIEW[Clinician evidence and edit screen]
    REVIEW --> APPROVE{Approve exact revision?}
    APPROVE -->|Edit| EDIT[New draft revision]
    EDIT --> REVIEW
    APPROVE -->|Yes| OUTBOX[(Approved submission intent)]
    OUTBOX --> ADAPT[EHR adapter and reconciliation]
    ADAPT --> EHR[(Epic or Oracle Health record)]
    STORE[(Restricted source and audit storage)] --- ALIGN
    STORE --- OUTBOX

The stream protocol carries a session ID, monotonic chunk sequence and time range. A reconnect resumes from the highest durably acknowledged chunk. Duplicate chunks must not duplicate transcript content; sequence gaps remain visible. If storage is disabled by the approved retention policy, define the shorter replay window and disclose that recovery limitation.

The client may keep a bounded encrypted retry buffer where device policy permits. Do not quietly continue ambient capture when the session ends, the clinician changes patients or authorization expires. Stop and require a new confirmed encounter session.

APIs and record contracts

Operation Contract
POST /encounters/{id}/dictation-sessions Verify clinician access and patient/encounter association; return scoped streaming capability
Audio stream Bounded frames, sequence IDs, acknowledgments, gap and finalization events
PATCH /notes/{id} Require expected draft revision and clinician authorization
POST /notes/{id}/approvals Bind approval to note hash, encounter and transcript/evidence revision
POST /notes/{id}/submissions Reserve one approved submission intent and return delivery state
Record Fields that prevent ambiguity
Encounter session Tenant, patient, encounter, clinician, recording scope, start/stop events
Audio chunk Session, sequence, time interval, checksum, retention state and access policy
Transcript segment Stable segment ID, revision, word timings, provisional/final state, speaker label and gaps
Clinical fact Source span, reporter, subject, negation, temporal context, value/unit and ambiguity
Note revision Text/template version, evidence references, author edits, unresolved checks and hash
Submission Approved hash, EHR target, correlation ID, request status, external identifier and reconciliation history

The model sees only data for the confirmed encounter. Necessary chart context should be explicitly labeled as chart-derived, with its record version and time; do not present it as something spoken in the current visit. Use access controls on every read and write.

Preserve meaning before polishing prose

Consider the synthetic utterance: “The patient reports a headache for three days and says they took acetaminophen 500 milligrams twice.” It does not establish a daily frequency, current prescription or a newly ordered medication.

Architecture / visual model
flowchart LR S[Utterance and audio timestamp] --> E[Entities and relations] E --> SYM[Reported headache for three days] E --> MED[Reported acetaminophen 500 mg] E --> FREQ[Two occasions reported with schedule unclear] E --> ABS[No vitals stated] SYM --> NOTE[Evidence-linked draft] MED --> NOTE FREQ --> ASK[Clinician clarification] ABS --> NOTE
Read diagram source
flowchart LR
    S[Utterance and audio timestamp] --> E[Entities and relations]
    E --> SYM[Reported headache for three days]
    E --> MED[Reported acetaminophen 500 mg]
    E --> FREQ[Two occasions reported with schedule unclear]
    E --> ABS[No vitals stated]
    SYM --> NOTE[Evidence-linked draft]
    MED --> NOTE
    FREQ --> ASK[Clinician clarification]
    ABS --> NOTE
Source wording Faithful representation Unsafe transformation
“No chest pain” Negated symptom, with speaker and time “Chest pain”
“My mother has diabetes” Family history Patient diagnosis
“I used to take it” Historical use with timing unresolved Current medication
“500 milligrams twice” Dose reported; schedule unclear A twice-daily prescription
“Fifteen—sorry, fifty” Linked correction requiring context confirmation Always choosing the last number regardless of speaker
No vital sign stated Not documented Normal vital signs

A terminology dictionary can help recognize medication names and abbreviations. It cannot safely expand every abbreviation without context. A learned NER model is not inherently deterministic or error-free; even deterministic execution can produce consistently wrong extraction.

Diarization should preserve “unconfirmed” roles until mapped through a reliable workflow. A visitor can sound similar to the patient, and speakers may overlap. Avoid enrollment voiceprints by default; if required, treat biometric identity, permissions and retention as a separate design decision. Noise reduction can remove useful speech as well as noise, so evaluate the processed audio against the original.

Handle transcript revisions explicitly

Partial ASR text may change as more audio arrives. Draft only from a pinned finalized transcript revision and invalidate generated sections when their source changes. Keep a clinician's manual edit as an edit, rather than overwriting it with a late model response. Show a comparison when source changes affect an edited section.

The following function is a pure review gate over validated server-side records. It does not perform the EHR write or replace role authorization. Approvals include both the note hash and the supporting transcript revision.

def note_submission_state(note, session):
    if (note["patient_id"], note["encounter_id"]) != (
        session["patient_id"], session["encounter_id"]
    ):
        return "IDENTITY_MISMATCH"
    if note["transcript_revision"] != session["transcript_revision"]:
        return "SOURCE_CHANGED"
    if note["unresolved_checks"]:
        return "CLINICIAN_REVIEW_REQUIRED"
    approval = note.get("approval")
    if not approval or approval["note_hash"] != note["content_hash"]:
        return "APPROVAL_REQUIRED"
    if approval["transcript_revision"] != note["transcript_revision"]:
        return "APPROVAL_REQUIRED"
    return "ELIGIBLE_FOR_AUTHORIZED_SUBMISSION"

Run these checks in the same transactional revision check that reserves the submission. Otherwise the note could change between validation and enqueueing. The EHR worker sends the reserved immutable bytes. If audio is incomplete, a clinician may complete the documentation from their own knowledge, but the record must distinguish those edits and their authorship from captured evidence.

EHR integration is a workflow contract

The FHIR R4 example below is a generic draft document reference, not a complete vendor-specific request. status: current describes the reference's status; docStatus: preliminary describes the underlying document. Neither field creates a clinician's signature. The Base64 content is only “SOAP note.” The context.encounter field is an array. FHIR R4 DocumentReference.

{
  "resourceType": "DocumentReference",
  "status": "current",
  "docStatus": "preliminary",
  "type": {"text": "Nursing encounter note"},
  "subject": {"reference": "Patient/example-patient"},
  "author": [{"reference": "Practitioner/example-clinician"}],
  "content": [{"attachment": {
    "contentType": "text/plain",
    "data": "U09BUCBub3Rl"
  }}],
  "context": {
    "encounter": [{"reference": "Encounter/example-encounter"}]
  }
}

Confirm the hospital's capability statement, licensed/enabled endpoint, OAuth access, accepted terminology, author attribution and signing policy. Epic documents a clinical-note create interaction. Oracle Health Millennium's create documentation permits final for provider access and final or amended for system access. Therefore, do not send this generic preliminary example to that endpoint or turn a draft into final merely to satisfy its schema. Keep drafts in the review workspace until the configured authorization and filing contract is satisfied. Epic clinical-note creation, Oracle Health creation contract.

A note mentioning a plan is not a medication order. Order entry, diagnosis coding and structured medication reconciliation require separate authorized workflows and evaluation; a document upload must not silently create them.

Architecture / visual model
stateDiagram-v2 [*] --> Draft Draft --> Reviewed: Clinician resolves gaps Reviewed --> Draft: Note or source changes Reviewed --> Approved: Exact revision approved Approved --> Draft: Note or source changes before reservation Approved --> Pending: Authorized submission reserved Pending --> Confirmed: External record verified Pending --> Unknown: Timeout after dispatch Unknown --> Confirmed: Reconciliation finds matching record Unknown --> Pending: Proven absent and safe retry Confirmed --> Amendment: Clinician initiates correction
Read diagram source
stateDiagram-v2
    [*] --> Draft
    Draft --> Reviewed: Clinician resolves gaps
    Reviewed --> Draft: Note or source changes
    Reviewed --> Approved: Exact revision approved
    Approved --> Draft: Note or source changes before reservation
    Approved --> Pending: Authorized submission reserved
    Pending --> Confirmed: External record verified
    Pending --> Unknown: Timeout after dispatch
    Unknown --> Confirmed: Reconciliation finds matching record
    Unknown --> Pending: Proven absent and safe retry
    Confirmed --> Amendment: Clinician initiates correction

Store the submission identifier before calling the EHR. Prefer supported conditional create or vendor idempotency semantics, but verify support rather than assuming the FHIR standard enables it everywhere. If the result is uncertain, query by supported correlation fields or route to manual reconciliation. Blind retry can create duplicate notes. A confirmed write is not the same as a signed note unless the configured workflow explicitly establishes that state.

Privacy and deployment choices

HIPAA does not require on-premises processing. HHS permits cloud processing with the appropriate business associate agreement (BAA) and compliance safeguards. A BAA is a contractual requirement, not evidence that every configuration, endpoint or downstream processor is approved. Encryption alone is insufficient; even an encrypted-data cloud provider without the key can have business-associate obligations. HHS cloud guidance.

Option Benefits Responsibilities and costs
Approved cloud ASR and note service Elastic capacity and managed serving Endpoint eligibility, contracts, access, retention, geography and outage plan
Local ASR and note serving Direct operational control over processing GPU capacity, failover, patching, model evaluation and staff time
Hybrid deployment May keep selected data or processing local Every boundary still needs review; sending transcripts externally still discloses patient data

A current streaming candidate is GPT-Live-Transcribe; a local baseline is Whisper. Neither reference establishes clinical suitability, a particular latency or healthcare-contract eligibility. Choose the qualified deployment after representative evaluation.

Define recording permission, pause/stop behavior, retention and deletion separately for audio, drafts, final records, backups and troubleshooting traces. Do not assume one universal consent or retention rule. Avoid raw PHI in metrics and generic application logs. Restrict evidence replay to authorized clinicians and record access. Urgent clinical concerns follow the hospital's existing escalation process; this documentation assistant is not a monitored emergency channel.

Failure tests and quality metrics

Test case Expected behavior
Similar patient names or chart switch during capture Stop or reject mismatched context; never infer the destination from speech
Audio dropout during a medication statement Expose the gap and require clinician completion
Speaker overlap or family member history Preserve uncertain attribution; do not convert it to a patient fact
“No” or a decimal point disappears Detect or surface source discrepancy for review
Final transcript revises an already drafted sentence Invalidate affected generated content and preserve manual edits
Model returns an instruction embedded in speech Treat transcript as data; no tools or orders are authorized by it
EHR times out after accepting a note Reconcile the existing submission before retrying

Measure word error rate, but also clinically significant substitution, omission and hallucination rates. A transcript with low overall word error can still contain a dangerous dosage error. Evaluate source-to-transcript and transcript-to-note separately: comparing the note only with extracted entities misses errors that extraction already omitted.

Use clinician-adjudicated recordings across wards, accents, languages, noise conditions, terminology and speaker combinations. Report sample sizes and confidence intervals; split evaluation by encounter and speaker to reduce leakage. Assess retained meaning after edits, not just fluent note text. Measure clinician correction time, rejected drafts, chart mismatch incidents, duplicate writes and confirmed filing latency.

Costs and decision tradeoffs

For 150,000 audio minutes, assume a budget allowance of $0.01/minute, giving $1,500/month. This is not an advertised provider rate. Include retranscription and overlapping windows in measured billable minutes.

A sample text-drafting budget uses GPT-6 Sol at $2 input and $10 output per million tokens. At 4,000 input and 800 output tokens, one pass costs $0.016. Two passes for 10,000 notes cost $320/month. This is a pricing example, not a clinical model recommendation; choose an eligible, evaluated deployment. OpenAI pricing.

Item Monthly planning cost
Transcription allowance $1,500
Two text passes per encounter $320
Storage, integration workers and monitoring allowance $1,000
Automation subtotal $2,820
Two review minutes × 10,000 notes × $60/hour $20,000
Partial operating total $22,820, or $2.282/encounter

The review duration is an assumption to validate, not a clinical recommendation. Exclude neither correction time nor operational staffing when comparing cloud with local serving. A reduction from five to two documentation minutes saves 500 hours/month at this volume; that benefit must be demonstrated without degrading documentation quality.

Decision Benefit Tradeoff
Dictation first Lower attribution complexity and deliberate recording Does not capture an entire conversation automatically
Provisional transcript feedback Clinician notices capture problems early Clearly label revisions; avoid constant draft rewriting
Evidence-linked facts Faster correction and more faithful notes More storage and careful source/version management
Draft-only model authority Keeps clinical and signing responsibility explicit Clinician review remains a real capacity requirement
Outbox plus reconciliation Prevents lost or casually duplicated filing Requires vendor-specific recovery behavior

Interview follow-ups

1. Why is diarization insufficient to identify the patient?

It clusters speaker audio; it does not verify the speaker's role or whom a statement describes. Confirm participant roles and preserve family-history and interpreter context. Dictation may avoid diarization entirely.

2. Can the latest statement always replace an earlier one?

No. An explicit correction by the same speaker about the same fact differs from a second person's conflicting account. Preserve both evidence spans and ask the clinician to resolve ambiguity.

3. Why not write notes to the EHR as they stream?

Partial transcripts change, the note may contain unsupported facts and the endpoint may file a final document. Keep revisions in the review workspace, then submit only the authorized immutable version through the verified workflow.

4. How do we test medical terminology?

Use adjudicated clinical audio with medication names, amounts, units, negation and abbreviations. Test hints and noise reduction against that data. A dictionary or high model confidence is not proof of correct meaning.

5. What happens when the cloud model is unavailable?

Keep manual documentation available. An approved local fallback is optional and must have its own quality and capacity evidence. Show pending audio and drafts without pretending they have been transcribed or filed.

6. Does human review eliminate risk?

No. Reviewers may miss errors or over-trust polished drafts. Highlight uncertain evidence, preserve source playback where allowed and measure post-review accuracy and correction effort. Avoid interfaces that encourage approving unseen text.

Closing notes

A strong design preserves identity, meaning and approval across a changing audio stream and a stateful clinical record. Start with scoped dictation, add evidence-linked drafting, then earn ambient recording through evaluation. Close the interview by naming the hardest remaining problems: source omissions, role attribution, clinician workload and each EHR's exact filing semantics.

Related: Realtime voice agents, Reliability patterns.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Case Study: Pharmaceutical Promotion Review
NEXT LESSONCase Study: Real-Time Payment Fraud Decisions →

Explore the diagram