This is a hypothetical interview scenario. Workload, latency, staffing and costs are planning assumptions, not measured results. Model and interoperability facts have primary-source links. Clinical review and hospital-approved workflows remain part of the design.
Interview focus: produce faithful draft notes from speech, preserve patient and encounter identity, and file only an authorized version through the EHR's actual interface.
60-second interview answer
I would begin with clinician-initiated dictation for a confirmed patient encounter, a versioned transcript and an editable draft note. Streaming transcription gives immediate feedback, while stable segments feed clinical extraction and note generation. Every proposed fact retains its speaker, time, negation and source. Clinicians resolve uncertain content and approve the exact note before filing. The EHR adapter verifies permissions and reconciles uncertain writes. I would expand to ambient multi-speaker recording only after measuring attribution and omission errors. Success is accurate documentation with less correction time, including privacy, review and integration costs.
Remember: Confirm encounter → Capture → Preserve meaning → Review → File exact note.
Interview problem and scope
A hospital network wants nurses to document encounters by voice instead of retyping notes. Design a service for 10,000 encounters per month that supports noisy rooms, medical terminology and integration with Epic and Oracle Health.
First distinguish two products. Dictation records a clinician intentionally describing a visit. Ambient documentation records a conversation and must handle patients, clinicians, family members, interpreters and overlapping speech. Start with dictation; make ambient capture an explicit later requirement. Neither product should independently diagnose, prescribe or place orders.
Functional requirements
- Start a recording only after the authenticated clinician confirms the patient, encounter and required recording permissions.
- Stream provisional text, preserve finalized segment revisions and visibly report capture gaps.
- Extract stated clinical facts with supporting transcript/audio locations and unresolved ambiguity.
- Draft the hospital's approved note template; support edits, comparison with source evidence and clinician approval.
- File an approved version through the configured EHR integration and show pending, confirmed or failed delivery.
- Support corrections and amendments without silently rewriting a previously signed clinical record.
Nonfunctional requirements
- Meaning fidelity: evaluate medication, allergy, negation, temporal and speaker errors independently of overall transcription accuracy.
- Responsiveness: target first provisional text under 500 ms p95 from the agreed available-speech point; measure finalization and note generation separately.
- Identity safety: bind all artifacts and writes to the server-verified patient and encounter; a spoken name cannot switch charts.
- Privacy: approve every processor of protected health information, including audio storage, ASR, note models, support tooling and logs.
- Resilience: keep the existing manual documentation workflow available during model or EHR outages; never fabricate missing audio.
- Auditability: record source revisions, model/configuration versions, clinician edits, approvals and external write receipts.
Definitions to establish before the diagram
| Term | Standard meaning | Boundary in this design |
|---|---|---|
| ASR | Automatic speech recognition converts speech to text | It can mishear words, numbers and negation |
| VAD | Voice activity detection identifies speech intervals | It does not identify a speaker or determine clinical importance |
| Diarization | Separates audio by speaker identity labels | “Speaker 2” does not establish patient or clinician role |
| NER | Named entity recognition locates entities such as a drug or symptom | Relations, certainty, temporality and negation require additional context |
| EHR | Electronic health record | The destination has its own permissions, templates and signing workflow |
| FHIR | HL7's standard for exchanging healthcare information | Standard-shaped JSON does not guarantee an enabled write API |
SOAP is one possible template: Subjective contains reported symptoms, Objective contains observations and measurements, Assessment contains the clinician's documented assessment, and Plan contains the clinician's stated next steps. Use the institution's nursing template when SOAP is inappropriate. A missing assessment is not an invitation for the model to invent one.
Capacity and latency estimates
Assume 15 recorded minutes per encounter and 20 eight-hour working days per month.
| Quantity | Calculation | Implication |
|---|---|---|
| Recorded minutes | 10,000 × 15 = 150,000/month | Bill transcription by actual processed audio, including retries |
| Stream hours | 150,000 / 60 = 2,500/month | Mean concurrency over 160 working hours is 15.6 streams |
| Peak concurrency | Five times mean ≈ 79 streams | Benchmark at peak plus reconnect and redundancy headroom |
| Raw mono audio | 16,000 samples/s × 2 bytes × 900 s × 10,000 | 288 GB/month before replicas and derived artifacts |
| Compressed audio example | 64 kilobits/s × 150,000 minutes | 72 GB/month in decimal units; codec support and quality must be validated |
A machine that transcribes one completed recording at ten times real time does not necessarily support ten concurrent streams at a tight tail-latency target. Measure streaming batching, memory, scheduling and failover behavior with representative clinical audio.
| Clock | Illustrative target | Start and end |
|---|---|---|
| First provisional text | Under 500 ms p95 | Enough speech available to the client → first displayed partial |
| Stable utterance | Under 1 second p95 | Actual end of utterance → final segment displayed |
| Draft update | Under 3 seconds p95 | Stable transcript revision → corresponding draft ready |
| Approved note | Measured, no model-only promise | Encounter documentation start → clinician approval |
| EHR confirmation | Contract-specific | Approved submission → verified external record |
A first-partial planning budget might allocate 50 ms to framing, 60 ms to transport, 80 ms to queueing, 200 ms to recognition and 50 ms to rendering: 440 ms total. These component allowances are not measured percentiles; adding individual p95 values does not produce an end-to-end p95. Endpointing—the decision that a speaker has finished—adds a separate delay.
Baseline, then detailed design
The baseline is push-to-talk dictation, an approved transcription engine and a manually edited note. Bind the recording to the selected encounter, retain revision history and require normal clinician approval. This yields value without introducing multi-speaker role inference or autonomous clinical reasoning.
Add structured extraction, evidence-linked drafting and ambient recording only when they improve measured documentation effort and quality.
Read diagram source
flowchart TB
UI[Clinician confirms patient encounter and recording] --> CAP[Capture with pause indicator and sequence numbers]
CAP --> STREAM[Authenticated audio stream]
STREAM --> ASR[Approved streaming ASR]
STREAM --> SPEAK[Optional audio diarization]
ASR --> ALIGN[Versioned words and speaker alignment]
SPEAK --> ALIGN
ALIGN --> FACT[Clinical facts with evidence and uncertainty]
ALIGN --> DRAFT[Template-aware note draft]
FACT --> DRAFT
DRAFT --> CHECK[Source fidelity and missing-content checks]
CHECK --> REVIEW[Clinician evidence and edit screen]
REVIEW --> APPROVE{Approve exact revision?}
APPROVE -->|Edit| EDIT[New draft revision]
EDIT --> REVIEW
APPROVE -->|Yes| OUTBOX[(Approved submission intent)]
OUTBOX --> ADAPT[EHR adapter and reconciliation]
ADAPT --> EHR[(Epic or Oracle Health record)]
STORE[(Restricted source and audit storage)] --- ALIGN
STORE --- OUTBOX
The stream protocol carries a session ID, monotonic chunk sequence and time range. A reconnect resumes from the highest durably acknowledged chunk. Duplicate chunks must not duplicate transcript content; sequence gaps remain visible. If storage is disabled by the approved retention policy, define the shorter replay window and disclose that recovery limitation.
The client may keep a bounded encrypted retry buffer where device policy permits. Do not quietly continue ambient capture when the session ends, the clinician changes patients or authorization expires. Stop and require a new confirmed encounter session.
APIs and record contracts
| Operation | Contract |
|---|---|
POST /encounters/{id}/dictation-sessions |
Verify clinician access and patient/encounter association; return scoped streaming capability |
| Audio stream | Bounded frames, sequence IDs, acknowledgments, gap and finalization events |
PATCH /notes/{id} |
Require expected draft revision and clinician authorization |
POST /notes/{id}/approvals |
Bind approval to note hash, encounter and transcript/evidence revision |
POST /notes/{id}/submissions |
Reserve one approved submission intent and return delivery state |
| Record | Fields that prevent ambiguity |
|---|---|
| Encounter session | Tenant, patient, encounter, clinician, recording scope, start/stop events |
| Audio chunk | Session, sequence, time interval, checksum, retention state and access policy |
| Transcript segment | Stable segment ID, revision, word timings, provisional/final state, speaker label and gaps |
| Clinical fact | Source span, reporter, subject, negation, temporal context, value/unit and ambiguity |
| Note revision | Text/template version, evidence references, author edits, unresolved checks and hash |
| Submission | Approved hash, EHR target, correlation ID, request status, external identifier and reconciliation history |
The model sees only data for the confirmed encounter. Necessary chart context should be explicitly labeled as chart-derived, with its record version and time; do not present it as something spoken in the current visit. Use access controls on every read and write.
Preserve meaning before polishing prose
Consider the synthetic utterance: “The patient reports a headache for three days and says they took acetaminophen 500 milligrams twice.” It does not establish a daily frequency, current prescription or a newly ordered medication.
Read diagram source
flowchart LR
S[Utterance and audio timestamp] --> E[Entities and relations]
E --> SYM[Reported headache for three days]
E --> MED[Reported acetaminophen 500 mg]
E --> FREQ[Two occasions reported with schedule unclear]
E --> ABS[No vitals stated]
SYM --> NOTE[Evidence-linked draft]
MED --> NOTE
FREQ --> ASK[Clinician clarification]
ABS --> NOTE
| Source wording | Faithful representation | Unsafe transformation |
|---|---|---|
| “No chest pain” | Negated symptom, with speaker and time | “Chest pain” |
| “My mother has diabetes” | Family history | Patient diagnosis |
| “I used to take it” | Historical use with timing unresolved | Current medication |
| “500 milligrams twice” | Dose reported; schedule unclear | A twice-daily prescription |
| “Fifteen—sorry, fifty” | Linked correction requiring context confirmation | Always choosing the last number regardless of speaker |
| No vital sign stated | Not documented | Normal vital signs |
A terminology dictionary can help recognize medication names and abbreviations. It cannot safely expand every abbreviation without context. A learned NER model is not inherently deterministic or error-free; even deterministic execution can produce consistently wrong extraction.
Diarization should preserve “unconfirmed” roles until mapped through a reliable workflow. A visitor can sound similar to the patient, and speakers may overlap. Avoid enrollment voiceprints by default; if required, treat biometric identity, permissions and retention as a separate design decision. Noise reduction can remove useful speech as well as noise, so evaluate the processed audio against the original.
Handle transcript revisions explicitly
Partial ASR text may change as more audio arrives. Draft only from a pinned finalized transcript revision and invalidate generated sections when their source changes. Keep a clinician's manual edit as an edit, rather than overwriting it with a late model response. Show a comparison when source changes affect an edited section.
The following function is a pure review gate over validated server-side records. It does not perform the EHR write or replace role authorization. Approvals include both the note hash and the supporting transcript revision.
def note_submission_state(note, session):
if (note["patient_id"], note["encounter_id"]) != (
session["patient_id"], session["encounter_id"]
):
return "IDENTITY_MISMATCH"
if note["transcript_revision"] != session["transcript_revision"]:
return "SOURCE_CHANGED"
if note["unresolved_checks"]:
return "CLINICIAN_REVIEW_REQUIRED"
approval = note.get("approval")
if not approval or approval["note_hash"] != note["content_hash"]:
return "APPROVAL_REQUIRED"
if approval["transcript_revision"] != note["transcript_revision"]:
return "APPROVAL_REQUIRED"
return "ELIGIBLE_FOR_AUTHORIZED_SUBMISSION"
Run these checks in the same transactional revision check that reserves the submission. Otherwise the note could change between validation and enqueueing. The EHR worker sends the reserved immutable bytes. If audio is incomplete, a clinician may complete the documentation from their own knowledge, but the record must distinguish those edits and their authorship from captured evidence.
EHR integration is a workflow contract
The FHIR R4 example below is a generic draft document reference, not a complete vendor-specific request. status: current describes the reference's status; docStatus: preliminary describes the underlying document. Neither field creates a clinician's signature. The Base64 content is only “SOAP note.” The context.encounter field is an array. FHIR R4 DocumentReference.
{
"resourceType": "DocumentReference",
"status": "current",
"docStatus": "preliminary",
"type": {"text": "Nursing encounter note"},
"subject": {"reference": "Patient/example-patient"},
"author": [{"reference": "Practitioner/example-clinician"}],
"content": [{"attachment": {
"contentType": "text/plain",
"data": "U09BUCBub3Rl"
}}],
"context": {
"encounter": [{"reference": "Encounter/example-encounter"}]
}
}
Confirm the hospital's capability statement, licensed/enabled endpoint, OAuth access, accepted terminology, author attribution and signing policy. Epic documents a clinical-note create interaction. Oracle Health Millennium's create documentation permits final for provider access and final or amended for system access. Therefore, do not send this generic preliminary example to that endpoint or turn a draft into final merely to satisfy its schema. Keep drafts in the review workspace until the configured authorization and filing contract is satisfied. Epic clinical-note creation, Oracle Health creation contract.
A note mentioning a plan is not a medication order. Order entry, diagnosis coding and structured medication reconciliation require separate authorized workflows and evaluation; a document upload must not silently create them.
Read diagram source
stateDiagram-v2
[*] --> Draft
Draft --> Reviewed: Clinician resolves gaps
Reviewed --> Draft: Note or source changes
Reviewed --> Approved: Exact revision approved
Approved --> Draft: Note or source changes before reservation
Approved --> Pending: Authorized submission reserved
Pending --> Confirmed: External record verified
Pending --> Unknown: Timeout after dispatch
Unknown --> Confirmed: Reconciliation finds matching record
Unknown --> Pending: Proven absent and safe retry
Confirmed --> Amendment: Clinician initiates correction
Store the submission identifier before calling the EHR. Prefer supported conditional create or vendor idempotency semantics, but verify support rather than assuming the FHIR standard enables it everywhere. If the result is uncertain, query by supported correlation fields or route to manual reconciliation. Blind retry can create duplicate notes. A confirmed write is not the same as a signed note unless the configured workflow explicitly establishes that state.
Privacy and deployment choices
HIPAA does not require on-premises processing. HHS permits cloud processing with the appropriate business associate agreement (BAA) and compliance safeguards. A BAA is a contractual requirement, not evidence that every configuration, endpoint or downstream processor is approved. Encryption alone is insufficient; even an encrypted-data cloud provider without the key can have business-associate obligations. HHS cloud guidance.
| Option | Benefits | Responsibilities and costs |
|---|---|---|
| Approved cloud ASR and note service | Elastic capacity and managed serving | Endpoint eligibility, contracts, access, retention, geography and outage plan |
| Local ASR and note serving | Direct operational control over processing | GPU capacity, failover, patching, model evaluation and staff time |
| Hybrid deployment | May keep selected data or processing local | Every boundary still needs review; sending transcripts externally still discloses patient data |
A current streaming candidate is GPT-Live-Transcribe; a local baseline is Whisper. Neither reference establishes clinical suitability, a particular latency or healthcare-contract eligibility. Choose the qualified deployment after representative evaluation.
Define recording permission, pause/stop behavior, retention and deletion separately for audio, drafts, final records, backups and troubleshooting traces. Do not assume one universal consent or retention rule. Avoid raw PHI in metrics and generic application logs. Restrict evidence replay to authorized clinicians and record access. Urgent clinical concerns follow the hospital's existing escalation process; this documentation assistant is not a monitored emergency channel.
Failure tests and quality metrics
| Test case | Expected behavior |
|---|---|
| Similar patient names or chart switch during capture | Stop or reject mismatched context; never infer the destination from speech |
| Audio dropout during a medication statement | Expose the gap and require clinician completion |
| Speaker overlap or family member history | Preserve uncertain attribution; do not convert it to a patient fact |
| “No” or a decimal point disappears | Detect or surface source discrepancy for review |
| Final transcript revises an already drafted sentence | Invalidate affected generated content and preserve manual edits |
| Model returns an instruction embedded in speech | Treat transcript as data; no tools or orders are authorized by it |
| EHR times out after accepting a note | Reconcile the existing submission before retrying |
Measure word error rate, but also clinically significant substitution, omission and hallucination rates. A transcript with low overall word error can still contain a dangerous dosage error. Evaluate source-to-transcript and transcript-to-note separately: comparing the note only with extracted entities misses errors that extraction already omitted.
Use clinician-adjudicated recordings across wards, accents, languages, noise conditions, terminology and speaker combinations. Report sample sizes and confidence intervals; split evaluation by encounter and speaker to reduce leakage. Assess retained meaning after edits, not just fluent note text. Measure clinician correction time, rejected drafts, chart mismatch incidents, duplicate writes and confirmed filing latency.
Costs and decision tradeoffs
For 150,000 audio minutes, assume a budget allowance of $0.01/minute, giving $1,500/month. This is not an advertised provider rate. Include retranscription and overlapping windows in measured billable minutes.
A sample text-drafting budget uses GPT-6 Sol at $2 input and $10 output per million tokens. At 4,000 input and 800 output tokens, one pass costs $0.016. Two passes for 10,000 notes cost $320/month. This is a pricing example, not a clinical model recommendation; choose an eligible, evaluated deployment. OpenAI pricing.
| Item | Monthly planning cost |
|---|---|
| Transcription allowance | $1,500 |
| Two text passes per encounter | $320 |
| Storage, integration workers and monitoring allowance | $1,000 |
| Automation subtotal | $2,820 |
| Two review minutes × 10,000 notes × $60/hour | $20,000 |
| Partial operating total | $22,820, or $2.282/encounter |
The review duration is an assumption to validate, not a clinical recommendation. Exclude neither correction time nor operational staffing when comparing cloud with local serving. A reduction from five to two documentation minutes saves 500 hours/month at this volume; that benefit must be demonstrated without degrading documentation quality.
| Decision | Benefit | Tradeoff |
|---|---|---|
| Dictation first | Lower attribution complexity and deliberate recording | Does not capture an entire conversation automatically |
| Provisional transcript feedback | Clinician notices capture problems early | Clearly label revisions; avoid constant draft rewriting |
| Evidence-linked facts | Faster correction and more faithful notes | More storage and careful source/version management |
| Draft-only model authority | Keeps clinical and signing responsibility explicit | Clinician review remains a real capacity requirement |
| Outbox plus reconciliation | Prevents lost or casually duplicated filing | Requires vendor-specific recovery behavior |
Interview follow-ups
1. Why is diarization insufficient to identify the patient?
It clusters speaker audio; it does not verify the speaker's role or whom a statement describes. Confirm participant roles and preserve family-history and interpreter context. Dictation may avoid diarization entirely.
2. Can the latest statement always replace an earlier one?
No. An explicit correction by the same speaker about the same fact differs from a second person's conflicting account. Preserve both evidence spans and ask the clinician to resolve ambiguity.
3. Why not write notes to the EHR as they stream?
Partial transcripts change, the note may contain unsupported facts and the endpoint may file a final document. Keep revisions in the review workspace, then submit only the authorized immutable version through the verified workflow.
4. How do we test medical terminology?
Use adjudicated clinical audio with medication names, amounts, units, negation and abbreviations. Test hints and noise reduction against that data. A dictionary or high model confidence is not proof of correct meaning.
5. What happens when the cloud model is unavailable?
Keep manual documentation available. An approved local fallback is optional and must have its own quality and capacity evidence. Show pending audio and drafts without pretending they have been transcribed or filed.
6. Does human review eliminate risk?
No. Reviewers may miss errors or over-trust polished drafts. Highlight uncertain evidence, preserve source playback where allowed and measure post-review accuracy and correction effort. Avoid interfaces that encourage approving unseen text.
Closing notes
A strong design preserves identity, meaning and approval across a changing audio stream and a stateful clinical record. Start with scoped dictation, add evidence-linked drafting, then earn ambient recording through evaluation. Close the interview by naming the hardest remaining problems: source omissions, role attribution, clinician workload and each EHR's exact filing semantics.
Related: Realtime voice agents, Reliability patterns.