Interview problem: moderate text and media at social-platform scale while limiting harmful exposure, avoiding wrongful restrictions, and providing timely human review and appeals.
Moderation applies a platform's policy to content and behavior. A classifier estimates signals; a separate decision process chooses an action such as allowing, labeling, limiting distribution, age-gating, restricting or escalating. Similar words can appear in harassment, journalism or a quotation, so a keyword is not a complete policy judgment.
This is a hypothetical Learnastra interview scenario. Volumes, routing fractions, costs and objectives are assumptions to validate, not measured production results.
1. Define policy and the visible product behavior
Clarify media types, languages, live versus uploaded content, geographic scope, user age restrictions, and what remains visible while assessment is pending. Distinguish immediate containment from completing a review and from any external reporting process.
Functional requirements
- Accept versioned posts and media, validate formats, and create a durable assessment job.
- Apply approved known-media signals, specialized classifiers and contextual review where needed.
- Make policy-versioned decisions with evidence and a reason code.
- Enforce visibility/account actions idempotently and prevent stale decisions from overwriting newer ones.
- Route urgent risks directly to trained specialists from any stage.
- Support human decisions, user notifications and eligible appeals.
- Record changes, reversals and evidence for audit; curate feedback for evaluation and future training.
Nonfunctional requirements
- Plan for 10M posts/day and 50M daily active users; size media processing from bytes, duration and sampling, not just post count.
- Define a bounded initial-visibility deadline, such as an illustrative p95 300ms for ordinary text. Deeper assessment may be asynchronous under the agreed policy.
- Target critical review completion within 15 minutes, high-priority within one hour, standard within 24 hours and appeals within seven days, each at a stated percentile. These are scenario SLOs, not universal legal deadlines.
- Measure harmful exposure as a fraction of views, alongside category-specific precision, recall and false-positive rates.
- Protect restricted media, reporter identities, account data and reviewer access; minimize ordinary logs.
- Bound retries, model spend and queues, with reserved specialist capacity and defined outage behavior.
- Version content, policy, models and decisions so an appeal or incident can reconstruct what happened.
An exposure objective below 0.1%, recall above 99% and precision above 95% are proposed goals. Evaluate whether they are simultaneously feasible for each category and population; a single aggregate score cannot establish that.
Content categories and action policy
| Category | Initial policy decision | Further handling |
|---|---|---|
| Suspected child sexual abuse material (CSAM) | Restrict access and route to the approved child-safety process | Trained specialists and reporting/evidence procedures |
| Violence/gore | Assess severity, context and imminent danger | Imminent risk may require urgent containment |
| Hate speech | Assess protected target, context and threat | Serious cases to appropriate reviewers |
| Harassment | Examine targeting, repetition and threats | Combine content and behavior evidence |
| Spam | Evaluate content, account and network signals | Rate limits, distribution controls or review |
| Misinformation | Apply topic-specific policy and evidence checking | Toxicity is not a truth detector |
| Adult content | Apply the platform's age/access policy | Suspected child exploitation takes the specialist route |
In the US, NCMEC's CyberTipline is a reporting channel for suspected online child exploitation. The service's child-safety/legal owners define applicable reporting and retention procedures. A model score does not determine legal obligations.
2. Work the confusion matrix before selecting thresholds
A false positive flags acceptable content; a false negative misses a violation. Precision is TP / (TP + FP), recall is TP / (TP + FN), and false-positive rate is FP / (FP + TN).
For 10,000 posts with 1% violations, 90% recall and a 1% false-positive rate:
| Actual class | Flagged | Not flagged | Total |
|---|---|---|---|
| Violation | 90 true positives | 10 false negatives | 100 |
| Acceptable | 99 false positives | 9,801 true negatives | 9,900 |
| Total | 189 | 9,811 | 10,000 |
Precision is 90 / 189 ≈ 47.6%, despite the apparently small 1% false-positive rate. 1 − precision is the fraction of flags that are false; it is not the false-positive rate.
A post-level miss and a view-level exposure have different consequences. One missed viral post can account for many harmful views. Estimate exposure from appropriately sampled impressions with reviewed labels, while protecting user privacy and accounting for sampling weights. Appeals and flagged-only review cannot estimate all missed violations.
3. Baseline and the failures that justify more tiers
Start with durable intake, a policy classifier, a versioned decision store and a staffed exception queue. Decide whether content is held or temporarily visible by risk category. A model timeout is an incomplete assessment, not an “allow” result.
| Baseline limitation | Added mechanism | Benefit | Cost or new failure |
|---|---|---|---|
| Repeated known harmful media | Approved known-media matching | Fast recognition of known items | Does not detect every novel item; match governance still matters |
| Rare contextual cases | Bounded contextual model review | More evidence-sensitive classification | Latency, cost, prompt injection and reviewer errors |
| Video/audio blind spots | Media-specific analysis and coverage records | Detects signals absent from text | Sampling gaps, transcription/OCR errors and compute |
| High-risk case waits behind ordinary traffic | Direct specialist path and reserved capacity | Faster urgent handling | Capacity and staffing requirements |
| Wrongful restriction persists | Independent appeal/review workflow | Correction and accountability | Selected feedback and additional load |
| Retried old decision undoes an appeal | Version-checked enforcement | Preserves the current decision | Coordination with feed/cache services |
PhotoDNA compares image signatures against known-image signatures. It is not facial recognition or a general detector of all new abusive content. Match results require the approved operational process; do not claim zero false positives or infer universal certainty from a hash hit.
4. Detailed architecture and cascade arithmetic
Average ingestion is 10M / 86,400 ≈ 116 posts/second; an assumed 10× peak is about 1,160/s. A single video can require much more processing than a short text post, so provision separate bounded media workers.
For one planning scenario, 5%, 85%, 8% and 2% of all submitted posts are resolved at tiers 1–4. These are routing assumptions, not conditional percentages or observed quality results.
| Tier | Work | Items processed/day | Resolved/day |
|---|---|---|---|
| 1 | Known-media and inexpensive signals | 10,000,000 | 500,000 |
| 2 | Specialized text/media classifiers | 9,500,000 | 8,500,000 |
| 3 | Contextual model review | 1,000,000 | 800,000 |
| 4 | Human review | 200,000 | 200,000 |
Every item reaching tier 3 has already incurred tiers 1 and 2. Urgent bypasses, appeals, rechecks after edits and representative quality audits add work beyond this simplified base flow.
Read diagram source
flowchart TD
UP[Versioned content intake] --> STORE[(Restricted content store)]
UP --> JOB[Durable jobs and admission controls]
JOB --> T1[Tier 1 known-media and fast signals]
T1 -->|Needs classification| T2[Tier 2 specialized classifiers]
T2 -->|Needs context| T3[Tier 3 bounded contextual review]
T3 -->|Unresolved| HUMAN[Human queues by severity and expertise]
T1 --> POLICY[Versioned policy decision service]
T2 --> POLICY
T3 --> POLICY
JOB -->|Urgent signal| SPECIAL[Immediate containment and specialist queue]
T1 -->|Urgent signal| SPECIAL
T2 -->|Urgent signal| SPECIAL
T3 -->|Urgent signal| SPECIAL
SPECIAL --> HUMAN
HUMAN --> POLICY
POLICY --> DB[(Decision ledger and transactional outbox)]
DB --> ENF[Version-checked enforcement worker]
ENF --> FEED[Feed, search and media visibility controls]
DB --> NOTICE[Reason notice and appeal entry point]
NOTICE --> APPEAL[Appeal review with current evidence]
APPEAL --> HUMAN
DB -. selected cases and random audits .-> EVAL[Evaluation and curated labels]
The policy service combines signals and explicit rules; it is not just a majority vote. Keep a decision's classification, temporary visibility and final enforcement status separate. A failed enforcement write must not appear as a successfully removed post.
5. Define the contracts and evidence
POST /content/{id}/assessments
{content_version, submission_id} → assessment_id
POST /decisions/{id}/appeals
{appeal_request_id, reason, additional_evidence_ref?} → appeal_id
GET /content/{id}/moderation-status
→ content_version, current_decision_id, visibility, reason_code, appeal_status
All endpoints enforce the caller's role and object access. External reporters do not gain access to restricted content merely by naming its ID.
| Record | Essential fields |
|---|---|
| Content version | Immutable media/text reference, hash, author, creation/edit time |
| Assessment | Content version, policy/model versions, category signals, coverage and incomplete checks |
| Decision | Stable ID, content version, decision sequence, policy rule, evidence, actor and time |
| Action job | Decision ID, destination, expected/current version and delivery status |
| Review item | Lane, original escalation time, deadline, language/expertise and lease |
| Appeal | Challenged decision, submitted evidence, reviewer and outcome |
An edit creates a new content version and assessment. A verdict for version 4 cannot automatically certify version 5. Retention follows the service's defined legal and operational policy; a database version history is not a reason to keep sensitive media forever.
6. What each automated tier can actually decide
Tier 1: fast signals
Known-media matching, account limits and pattern rules identify candidates cheaply. Keywords can indicate a need for context rather than a violation. Keep original and derived representations linked; normalization is evidence processing, not permission to destroy distinctions in legitimate language.
Tier 2: specialist classifiers and modality coverage
Use detectors evaluated for the platform's categories and media. As a current example, omni-moderation-latest accepts text/images and the standalone endpoint is free, but it does not assess audio. Several categories—including hate, harassment and sexual/minors—are text-only. Unsupported image categories can return zero; that means unassessed, not safe. Inspect applied input types. Do not send known or suspected CSAM to this API; use the dedicated child-safety process. OpenAI moderation documentation.
A multimodal general-purpose model may help with memes and context. OCR, speech transcription and specialist vision models can still be useful. Sampling ten video frames does not prove that every frame is safe. Record which media segments and modalities were actually checked.
Tier 3: contextual review
Supply the applicable trusted policy, submitted content, relevant conversation context and prior signals. Request a bounded schema with violation, no_violation or uncertain, a policy-rule reference and evidence locations. Keep the model's short evidence explanation distinct from a guarantee of correctness.
Schema validity establishes structure, not truth. Refusal, timeout, missing output or unsupported modality leaves the assessment incomplete. A provider's safety refusal is not itself proof of a violation under the platform's different policy. Rolling model aliases can change score distributions; recalibrate thresholds against held-out labels after changes.
Perspective API's official notice says service ends after December 31, 2026, with new usage/quota requests closed after February 2026. It is a migration concern for existing integrations, not a suitable new long-term dependency.
Executable example: missing coverage cannot become an allow decision
This small gate consumes trusted coverage metadata and an evaluated proposal. It demonstrates routing; it does not implement the classifiers or the full policy engine.
def assessment_route(*, urgent, required_checks, completed_checks, proposal):
if urgent:
return "specialist_containment"
if not required_checks or not set(required_checks) <= set(completed_checks):
return "incomplete_review"
if proposal not in {"violation", "no_violation", "uncertain"}:
return "incomplete_review"
if proposal == "uncertain":
return "human_review"
return "policy_decision" # Neither the detector nor this gate applies an action.
assert assessment_route(urgent=True, required_checks={"text"},
completed_checks=set(), proposal=None) == "specialist_containment"
assert assessment_route(urgent=False, required_checks={"text", "image"},
completed_checks={"text"}, proposal="no_violation") == "incomplete_review"
assert assessment_route(urgent=False, required_checks={"text"},
completed_checks={"text"}, proposal="uncertain") == "human_review"
Coverage should name concrete checks, such as image-violence or text-harassment, rather than only the coarse modality names used in this toy example. A completed check also needs a valid result for the required policy/model version.
7. Human review and appeal lifecycle
The completion clock starts at first qualifying escalation and includes queue waiting and active review. Retries do not restart it. An appeal gets its own clock from submission. Immediate containment protects distribution while those later decisions remain pending.
Read diagram source
stateDiagram-v2
[*] --> Submitted
Submitted --> Tier1
Submitted --> CriticalQueue: urgent risk
Tier1 --> Restricted: approved known-match action
Tier1 --> Tier2: needs classification
Tier1 --> CriticalQueue: urgent risk
Tier2 --> Decided: policy decision
Tier2 --> Tier3: needs context
Tier2 --> CriticalQueue: urgent risk
Tier3 --> Decided: policy decision
Tier3 --> CriticalQueue: urgent risk
Tier3 --> HighQueue: serious unresolved case
Tier3 --> StandardQueue: other unresolved case
CriticalQueue --> HumanReview: specialist starts
HighQueue --> HumanReview: reviewer starts
StandardQueue --> HumanReview: reviewer starts
HumanReview --> Decided: decision recorded
Restricted --> Decided: restriction recorded
Decided --> Appealed: eligible appeal submitted
Appealed --> AppealQueue
AppealQueue --> HumanReview: appeal reviewer starts
Decided --> Closed: process complete
Closed --> [*]
Use durable queues with leases/heartbeats, bounded retry and a dead-letter path. Separate severity lanes and reserve critical capacity. Within each lane, consider deadline, reach and calibrated uncertainty, with aging so ordinary cases are not starved. Language expertise and reviewer wellbeing constrain capacity, not merely the number of available workers.
The review UI shows the content version, surrounding context, applicable policy, evidence and decision history. Blur sensitive media by default and restrict access. For selected audits, conceal the model recommendation initially to measure independent judgment and reduce anchoring. Reviewers can disagree, request more evidence or escalate.
Enforcement must not undo a successful appeal
Write the decision and its action job in one database transaction. A worker applies it using the stable decision ID and a compare-and-set condition on the current content/decision version. A duplicate delivery is harmless only if the destination enforces that identity or the update is idempotent.
If decision 12 restores a post after appeal, a delayed worker for decision 11 must not restrict it again. Checking the ledger and then doing an unguarded remote write leaves a race: the destination must enforce the expected version, or all writes must pass through an authoritative serialized executor. Reconcile feed, search and media caches so the decision reaches actual distribution surfaces.
8. Reviewer capacity and economics
At 2% escalation, 10M posts produce 200,000 human cases/day. If 500 moderators each complete an assumed 200 reviews/day, capacity is 100,000/day and the backlog grows by 100,000/day before appeals and audits. The base flow needs 1,000 moderators working that day at that productivity, plus coverage/headroom. Do not assume every sensitive case takes the same time or that increasing throughput preserves review quality.
Cost worksheet with one denominator
The following are hypothetical per-item processing allowances, not model-provider quotes. Every row uses the earlier daily traffic counts.
| Component | Calculation | Cost/day |
|---|---|---|
| Fast signals | 10M × $0.0001 | $1,000 |
| Specialized classification | 9.5M × $0.0005 | $4,750 |
| Contextual review | 1M × $0.004 | $4,000 |
| Human review | 200,000 × $0.50 | $100,000 |
| Subtotal | Sum of all processed stages | $109,750 |
That is about $0.010975/post before media extraction, storage, retries, appeals, audits, specialist work and operational overhead. The $0.50 review allowance is not a wage quote. With token-billed models, compute calls × (input_tokens × input_rate + billed_output_tokens × output_rate) / 1M; include reasoning output, images/audio and context tiers under the actual provider contract.
A free detector endpoint still has quotas and integration costs. It reduces spending only where its category coverage and measured quality satisfy the requirement. Reducing escalation from 2% to 1% would halve the base review volume, but is acceptable only if the changed decision policy maintains the required error and exposure outcomes.
9. Robustness, failure policy and learning
| Evasion or failure | Defense to evaluate | Limitation |
|---|---|---|
| Character substitution | Locale-aware normalization and variant features | Can erase legitimate meaning |
| Text inside images | OCR plus visual/contextual assessment | OCR mistakes and visual-only signals |
| Invisible characters | Preserve source and inspect normalized views | Some characters carry real linguistic meaning |
| Context manipulation | Relevant conversation/behavior context | More data and privacy obligations |
| Encoded content | Bounded decoding and resource limits | Arbitrary recursive decoding is unsafe/expensive |
| Adversarial images | Test realistic crops, overlays and compression | No finite test proves all future variants are covered |
| Prompt injection | Treat posts/OCR as data; keep policy/action authority external | Prompt instructions alone are insufficient |
| Provider outage | Category-specific hold/restrict/review policy | Neither blanket allow nor blanket remove fits every risk |
Keep derived text, OCR and transcripts linked to their source and coverage. Never place restricted media in ordinary logs or send it to a provider merely because a general multimodal API accepts that file type.
Curate human decisions into datasets after quality checks; do not instantly retrain on every reviewer click. Appeals are selected by who contests a decision. Combine them with representative samples of allowed and restricted content, including weighted sampling when estimating population exposure. Measure reviewer disagreement and policy ambiguity separately from model error.
10. Evaluation and staged rollout
- Define category labels, action rules, annotation guidelines and severity before fitting thresholds.
- Evaluate by language, dialect, region, media and category, including legitimate contextual uses.
- Measure precision/recall/FPR, exposure, appeal reversals and incomplete assessments with clear denominators.
- Load-test the initial visibility deadline and review queues under bursts and provider failure.
- Test edited posts, duplicate jobs, out-of-order enforcement and appeals racing with old actions.
- Shadow new models/policies without changing decisions, then canary an approved scope.
- Roll back the model/policy configuration when required, while preserving newer individual appeal outcomes and an audit trail.
Policy owners define permitted interventions; ML owners evaluate detectors; operations owns queue capacity, reviewer quality and wellbeing; platform engineers own enforcement consistency. A dashboard should show oldest queued case and missed deadlines, not just average review time.
Interview follow-ups
1. Why can 99% overall accuracy be useless? With 1% violations, an always-allow classifier reaches 99% accuracy while missing every violation. Inspect the confusion matrix and exposure consequences instead.
2. Why keep the decision service separate from the classifier? A probability/label is evidence. Policy determines the action, context requirements, temporary visibility and appeal process. That separation supports versioning, auditing and different rules without pretending every score is a final verdict.
3. Can a zero category score establish an image is safe? No. The category may not support image inputs, or the model may have missed it. Check modality coverage and the applicable detector's evaluation before interpreting the score.
4. How would you stop stale enforcement after an appeal? Store a new decision version and require the destination or serialized executor to reject writes from older versions. Idempotency alone prevents duplicate actions; it does not order conflicting decisions.
5. Can appeals replace random audits? No. They overrepresent users willing and able to appeal and do not expose all harmful items that were allowed. Use appropriately sampled and reviewed population data as well.
6. What changes when the human queue exceeds capacity? Protect urgent lanes, expose the backlog, adjust staffing/scope and investigate upstream failures. Any threshold change must be evaluated against harm and wrongful-restriction costs; hiding cases from the queue does not solve them.
60-second interview answer
I would separate content signals, policy decisions and visibility enforcement. A bounded cascade handles routine cases, while urgent risks go directly to specialists and incomplete assessments follow an explicit temporary policy. Versioned decisions and guarded writes prevent old jobs from undoing appeals. I would size reviewer capacity from arrivals and handling time, measure exposure and both kinds of error by population, and roll out changes through shadow evaluation and controlled traffic with a correction path.
Remember: Policy → Prevalence → Pipeline → People → Appeals.