This is a hypothetical interview scenario. Workload, performance, staffing and costs are planning assumptions, not measured results. Regulatory references describe the US scope of this example and must be checked for the asset's jurisdiction and publication date.
Interview focus: design an evidence-backed review workflow that catches missing information, preserves source authority and binds qualified approval to the exact published asset.
60-second interview answer
I would start with a versioned asset inbox, a product evidence library and a qualified reviewer queue. Automation would extract explicit and implied claims, locate supporting sources and flag possible issues, including omissions and the advertisement's overall presentation. Each review would record its scope, source versions and unprocessed content. A clean automated report would still require qualified approval. I would then add parallel media processing, impact analysis for regulatory changes and a release gate that rejects changed assets. Success means fewer missed serious issues and a shorter measured review cycle, including reviewer time.
Remember: Scope → Capture → Check claims and context → Review → Approve exact version.
Interview problem and scope
A pharmaceutical marketing team submits 500 assets a month. Its current review cycle takes approximately two weeks. Design a system that prepares cited findings and routes reviews so the team can target a two-business-day turnaround.
Clarify what “compliance automation” means before drawing an architecture. Here it means assisting the company's medical, legal and regulatory review, with authorized people making release decisions. It does not mean obtaining FDA approval of each advertisement.
Start with US promotion of human prescription drugs by or on behalf of their manufacturers. Support print, web layouts and finished television/radio creative. Other countries, over-the-counter products and medical devices need separate applicability rules and expertise. Confirm whether draft storyboards or final rendered creative is being submitted; approving a script cannot approve a later video edit. FDA OPDP responsibilities.
Functional requirements
- Accept an immutable asset version with product, jurisdiction, audience, media, planned release date and accountable owner.
- Recover page text, layout, image regions, audio and timed video content, retaining locations and processing failures.
- Identify explicit claims, possible implied claims, omissions and presentation concerns against applicable sources.
- Present each finding with its location, supporting evidence, authority type, uncertainty and proposed reviewer routing.
- Support specialist review, revision, comments and approval for the exact asset and evidence versions.
- Identify active or pending assets affected by source changes and preserve a complete decision history.
Non-functional requirements
- Safety: measure recall for serious issues separately. No automated approval; failed media processing is visible and blocks completion.
- Turnaround: target 95% of complete submissions reaching a final company decision within two business days. Report time waiting for missing inputs separately without concealing total elapsed time.
- Automation latency: target a draft report within ten minutes for the agreed asset-size envelope; large videos use an asynchronous queue.
- Security: restrict unreleased product information by tenant, product and reviewer role; use approved data-processing arrangements.
- Auditability: retain original bytes, extracted evidence, source snapshots and revision-bound decisions under an approved retention schedule.
- Quality: specify the unit and denominator for every error target; measure reviewer burden as well as recall.
Interview tip: “Zero false negatives” expresses a safety ambition, not a statistically proven model capability. Define what a serious issue is, who adjudicates it and how unflagged material is sampled.
Size the human queue first
Assume 20 working days per month, six pages or media-equivalent segments per asset, 12 extracted claims per asset and a fivefold arrival burst.
| Quantity | Calculation | Design implication |
|---|---|---|
| Normal arrivals | 500 / 20 = 25 assets per working day | A durable job queue is sufficient |
| Burst day | 25 × 5 = 125 assets | Reserve review capacity and communicate queue age |
| Claims checked | 500 × 12 = 6,000 per month | Retrieval volume is modest; judgment dominates |
| Qualified review | 500 × 45 minutes = 375 hours/month | At 120 productive review hours/person, at least four reviewers before specialist bottlenecks |
| Video storage | 100 assets × 300 MB = 30 GB/month | Retain original plus derived media under explicit lifecycle rules |
Four reviewers provide 480 productive hours per month, approximately 32 average-length reviews per working day. A 125-asset burst takes nearly four days to clear even with no new arrivals. A two-day target needs surge capacity, fewer revisions or a scoped service commitment. Faster inference alone cannot solve it.
Begin with a useful baseline
Use object storage for versioned assets and sources, PostgreSQL for workflow records, deterministic metadata checks, text/layout extraction and a reviewer interface. A reviewer selects the applicable source set, checks claims and signs off. Full-text search may be sufficient for the initial evidence library.
This baseline already fixes lost attachments, uncertain approval versions and fragmented comments. Measure its review time and missed-issue rate before introducing an LLM. Add retrieval-augmented generation to draft evidence-backed findings where measured retrieval and reviewer productivity justify the complexity.
Detailed architecture
Read diagram source
flowchart TB
U[Authenticated asset owner] --> API[Submission API and immutable version]
API --> OBJ[(Original assets and derived media)]
API --> JOB[(Durable review jobs)]
JOB --> PARSE[Text layout audio and video workers]
PARSE --> COVER[Coverage manifest and extraction checks]
COVER --> CLAIM[Explicit and implied claim candidates]
CLAIM --> RET[Scoped hybrid evidence retrieval]
LIB[(Approved source editions and product evidence)] --> RET
RET --> ASSESS[Draft findings with exact citations]
COVER --> WHOLE[Whole-ad and missing-information checks]
ASSESS --> PACK[Review package and unresolved issues]
WHOLE --> PACK
PACK --> HUMAN[Medical legal and regulatory reviewers]
HUMAN --> GATE[Revision-bound release decision]
GATE --> CMS[Approved publishing integration]
DB[(Workflow and append-only decision events)] --- JOB
DB --- HUMAN
DB --- GATE
Workers claim bounded jobs with leases and retry transient failures. Each stage records the input hash and its output version. A changed asset starts a new review; reusing extraction is permissible only when the relevant bytes and processing configuration match. Deduplicate submission retries with a caller-scoped idempotency key.
APIs and records
| Operation | Important contract |
|---|---|
POST /assets/{id}/versions |
Validate metadata, finalize uploaded bytes and return a version plus content hash |
POST /reviews |
Pin asset version, review scope and approved source snapshot; return 202 and a job ID |
GET /reviews/{id} |
Return coverage, findings, evidence and workflow status to authorized roles |
POST /reviews/{id}/decisions |
Require current revision, authorized reviewer role and recorded rationale |
POST /releases |
Atomically verify approvals, unresolved blockers and exact content hash before reserving publication |
| Record | Required fields and purpose |
|---|---|
| Asset version | Content hash, storage key, media, language, audience, product, jurisdiction and intended release date |
| Source edition | Authority type, exact citation, official URL, content hash, publication date, effective interval and specialist approval |
| Claim or concern | Asset location, literal text or observed visual, claim category and extraction uncertainty |
| Finding | Applicable source passage, evidence relationship, severity, unresolved questions and reviewer disposition |
| Coverage manifest | Expected pages/tracks/segments, processed portions, supported checks, failures and manually resolved gaps |
| Decision | Reviewer identity/role, timestamp, approved artifact hash, source snapshot, rationale and superseded decision |
Store source text snapshots where permitted rather than depending on a URL whose content can change. A URL alone cannot reproduce a historical review. Source retrieval, asset access and decision APIs all enforce product access; hiding a button is insufficient.
Keep authority and applicability separate from similarity
Semantic similarity answers “is this passage related?” It does not establish legal authority, applicability or whether a claim is supported.
| Source | How it is used | Mistake to avoid |
|---|---|---|
| Statute or applicable regulation | Identify requirements within the legal scope and effective period | Treating a proposed amendment as current law |
| Final guidance | Understand the agency's recommendations and interpretation | Presenting guidance as interchangeable with a binding regulation |
| Approved product labeling | Establish product-specific indications, warnings and other relevant information | Using another strength, population or superseded label without review |
| Supporting study | Assess what the cited evidence actually measured and supports | Converting relative improvement into an absolute benefit |
| Warning or untitled letter | Examine a fact-specific enforcement example | Calling a nearest-neighbor match a binding precedent |
| Company policy | Apply additional internal standards | Describing an internal restriction as an FDA requirement |
The reviewer-approved source registry selects jurisdiction, product, audience, media and intended publication date before evidence ranking. Retrieve exact section identifiers and keyword matches alongside vector candidates. A finding must show how the source relates to the asset; a relevant citation does not establish that the conclusion follows from it.
For example, “80% of participants improved” and “symptoms improved by 80%” have different denominators and meanings. Check population, endpoint, comparator, duration and study limitations. If those facts are unavailable, show the missing evidence instead of manufacturing a regulatory quotation.
Fine-tuning may improve extraction or finding format after evaluation. It does not replace a current source library or supply an auditable citation from model memory.
Updating rules and enforcement examples
Read diagram source
flowchart LR
FEED[Official source change] --> DIFF[Version and classify change]
DIFF --> EXPERT[Specialist checks authority applicability and dates]
EXPERT --> FUTURE[Schedule approved future-effective edition]
EXPERT --> NOW[Publish applicable source snapshot]
FUTURE --> NOW
NOW --> IMPACT[Find dependent products rules and assets]
IMPACT --> QUEUE[Re-review or withdraw release eligibility]
NOW --> INDEX[(Search index with source lineage)]
Monitor official sources on a schedule with fetch failures and stale-source alerts. A monthly warning-letter import is useful background but cannot serve as the sole change detector. Preserve earlier editions for historical audits. A planned change may require preparation without changing today's release rule.
Maintain explicit source-to-review dependencies plus product/media metadata. Dependency links alone miss assets where the earlier retrieval failed to find the relevant rule; broader impact queries and specialist review cover that gap.
Evaluate the whole advertisement
A claim-by-claim checker misses omissions, font size, placement, visual implication and the relationship between spoken benefits and risks. A smiling or running person may suggest a benefit, but an image alone does not establish a violation.
For the in-scope consumer television/radio prescription-drug advertisements, the major-statement provisions address understandable language and audio, and television also requires concurrent audio/text, readable text and avoidance of distracting elements that impair comprehension. This calls for synchronized media evidence, not only a transcript. Apply the exact rule and exceptions through qualified review. 21 CFR 202.1(e), FDA final-rule questions and answers.
Read diagram source
flowchart LR
VIDEO[Final rendered creative] --> TRACKS[Audio transcript and time-aligned frames]
TRACKS --> TEXT[Claim and risk text with locations]
TRACKS --> VIEW[Readability duration contrast and audio checks]
TEXT --> TIMELINE[Combined evidence timeline]
VIEW --> TIMELINE
TIMELINE --> PERSON[Reviewer plays exact segment in full context]
PERSON --> DECISION[Finding or documented resolution]
Sampled frames can miss a short disclaimer or transition. Use shot changes and text-change detection to select intervals, disclose unsampled or undecodable spans, and let a reviewer inspect the original video. Automated readability and speech-rate metrics are screening aids, not a complete legal assessment. Unsupported languages or inaccessible audio produce a coverage gap, not a clean report.
Separate completion, risk and approval
Severity is potential impact; confidence is uncertainty about a finding. A low-confidence, potentially serious concern still needs appropriate attention. Neither is equivalent to review completion.
The following pure function assigns a queue state to an authenticated, server-created review record. The dictionaries represent validated application records. It intentionally cannot approve an asset.
def review_route(review):
if review["asset_hash"] != review["reviewed_asset_hash"]:
return "STALE_REVIEW"
if review["source_snapshot"] != review["required_source_snapshot"]:
return "SOURCE_RECHECK_REQUIRED"
if not review["coverage_complete"] or review["missing_evidence"]:
return "INCOMPLETE_REVIEW"
if any(f["severity"] == "high" and f["disposition"] == "unresolved"
for f in review["findings"]):
return "PRIORITY_SPECIALIST_REVIEW"
return "QUALIFIED_REVIEW_REQUIRED"
Coverage completion includes documented manual resolution of automated gaps, with reviewer identity and evidence; a client cannot simply set the Boolean. The release service separately requires all applicable role approvals and resolved blockers. Use a database transaction or compare-and-swap to verify the current review revision and reserve publication. The publisher consumes only the approved bytes.
If the publisher times out, reconcile its release identifier before retrying. An asset edited after approval must not inherit that approval. When a source change invalidates release eligibility, stop queued publication and evaluate already published assets through an explicit withdrawal/review process.
Quality and failure testing
Create an independently adjudicated evaluation set with clean assets, known issues, ambiguous cases, omissions, visual/audio concerns, languages and outdated-label traps. Hold out product families and later time periods to reduce leakage. Evaluate the complete pipeline, including parsing and retrieval.
| Metric | Worked example | Interpretation |
|---|---|---|
| Serious-issue recall | 190 of 200 adjudicated serious issues found = 95% | Ten misses still require analysis; do not blend them with minor issues |
| Precision | 190 supported flags / 240 flags = 79.2% | Approximately one in five flags creates avoidable review work |
| False-positive rate | 50 false positives / 800 negative units = 6.25% | Different denominator from the 20.8% false-discovery proportion |
| Coverage | Fully processed assets / submitted assets | A failed extractor cannot count as an issue-free asset |
| Review effort | Minutes reviewing findings, unflagged content and revisions | Avoid optimizing only time spent on the AI report |
| End-to-end turnaround | Submission to final disposition | Include queueing, specialist review and resubmission |
Use a clearly defined evaluation unit: an adjudicated issue opportunity, claim or whole asset. The numerical confusion example uses 1,000 labeled issue opportunities, of which 200 are positive. Do not combine claim-level recall with asset-level false positives in one confusion matrix.
Finding zero misses in 300 independent representative positive cases gives only an approximate 95% upper bound of 1% on the miss probability, using the rule of three. Dependence, distribution shift and insufficient coverage weaken that inference. Keep qualified review and monitor unflagged content.
| Failure introduced | Required behavior |
|---|---|
| OCR drops a page containing risk text | Mark incomplete coverage and expose the missing page |
| A retrieved warning letter concerns another product | Exclude it as direct support; optionally show labeled contextual evidence |
| Source update arrives after a draft finding | Re-evaluate affected dependencies and release eligibility |
| Ad contains instructions to ignore the risk statement | Treat them as asset content, never as agent instructions |
| Reviewer approves version 7 while version 8 is submitted | Preserve approval for version 7 only |
| No AI findings | Continue the normal qualified approval workflow |
Costs and tradeoffs
Use an initial allowance per asset; these are planning budgets rather than API quotes. Measure actual tokens, media processing and reviewer minutes before procurement.
| Stage | Allowance per asset | Why retain it |
|---|---|---|
| Parsing and media preparation | $0.25 | Preserves evidence locations and processing coverage |
| Claim extraction and evidence assessment | $1.80 | Helps reviewers find and compare support |
| Retrieval and report assembly | $0.12 | Produces a reproducible review package |
| Storage and orchestration allocation | $0.20 | Retains versions, jobs and audit events |
| Automation subtotal | $2.37 | 500 assets cost $1,185/month |
| Qualified review: 45 minutes at $120/hour | $90.00 | 500 assets cost $45,000/month |
| Partial operating total | $92.37 | $46,185/month, excluding integrations and source maintenance |
If an evaluated change reduces review from 45 to 35 minutes per asset without degrading quality, it saves 83.3 hours or $10,000 of modeled monthly labor value. That is capacity released, not necessarily a cash saving or permission to remove specialist approval. Add the cost of revisions and maintaining the regulatory corpus.
| Choice | Benefit | Cost or limitation |
|---|---|---|
| Full-text baseline before vector search | Easier exact-section retrieval and debugging | Lower recall for differently worded concerns |
| Hybrid retrieval | Combines exact identifiers and semantic candidates | Needs applicability filters, ranking evaluation and lineage |
| Conservative serious-issue routing | Reduces unattended high-impact ambiguity | More specialist load; calibrate with real review outcomes |
| Version-bound approvals | Prevents accidental approval reuse | New versions require deliberate re-review |
| Complete media evidence | Exposes presentation and omission concerns | More processing and human inspection than text alone |
Interview follow-ups
1. Can we automatically approve assets with no flags?
No, not in this design. Missing extraction, omitted risks and unknown applicability may create no flags. The report supports a qualified decision; coverage checks and normal approval still apply to unflagged assets.
2. How do we keep the model current?
Publish specialist-approved source editions with applicability and effective dates. Rebuild or incrementally update retrieval, then find affected pending and published assets. Model replacement alone does not update the legal basis of an earlier review.
3. Why not just retrieve similar warning letters?
They are useful fact-specific enforcement examples, but the applicable requirement and product evidence must support the actual finding. Similar language or a large count of letters does not prove a violation.
4. What does the audit explain?
The exact reviewed asset, evidence passages, source versions, supported checks, unresolved gaps, reviewer actions and final rationale. An evidence-based explanation is reproducible material; it is not a claim to expose the model's private reasoning.
5. What if a serious finding has low confidence?
Route it to a specialist with the uncertainty and missing evidence stated. Severity and confidence control different decisions. Lowering priority because the model is unsure can hide exactly the cases that require expert judgment.
6. How would you prove the two-day improvement?
Compare end-to-end review cohorts with similar products, media and complexity. Measure queue age, review minutes, revision counts and adjudicated quality, including unflagged assets. A ten-minute AI report is not evidence of a two-day approval process.
Closing notes
The core artifact is an approved asset version with a reproducible evidence package. Begin with reliable intake, source authority and human workflow. Add model assistance where it improves demonstrated issue detection or review effort. In an interview, close with the remaining risks: applicability judgment, media coverage, reviewer capacity and source changes after approval. Name an owner and an operational response for each.
Related: Document intelligence, Guardrails, Compliance and governance.