Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Design a Contract Document-Intelligence Pipeline

By Anup Rai11 min readReviewed September 2026

Interview problem: convert native and scanned contracts into searchable, structured records whose important fields can be traced to the exact source evidence.

Document intelligence combines text/layout recovery, information extraction and validation. OCR recognizes text in images. Extraction proposes field values from source material. Validation checks format, evidence and domain relationships. None of these automatically determines the correct legal interpretation of an ambiguous contract.

This is a hypothetical Learnastra interview scenario: 50,000 contracts/month, 2–200 pages each, in English, German, French and Spanish. The design must explain missing evidence, amendments, review capacity and cost per accepted record, not just produce plausible JSON.

1. Requirements and the output contract

Clarify the exact fields, whether signed amendments and exhibits are included, which source version governs a record, and which outputs are allowed to drive downstream actions. A search index and an automatic payment workflow need different acceptance policies.

Functional requirements

  1. Validate and store authorized PDFs with stable source/version identity.
  2. Recover native text or scan content at page level, preserving layout and evidence coordinates.
  3. Extract parties, dates, obligations, termination conditions and payment terms into typed records.
  4. Follow relevant definitions, cross-references, exhibits and amendments.
  5. Distinguish found, not found, ambiguous and processing-failed values.
  6. Route exceptions to a reviewer who can inspect the original source and correct the record.
  7. Publish approved records and searchable references with lineage and update history.

Nonfunctional requirements

  1. Process 50,000 documents/month with bounded queues, page/token limits and replayable jobs.
  2. Treat the proposed “95%+ key-field accuracy” as a metric to define: specify fields, matching rules, missing values and evaluation population.
  3. Propose p95 processing below ten minutes for a supported 100-page document, excluding separately reported human waiting time.
  4. Target under $0.50 average automated processing cost/document and separately report review and long-document costs.
  5. Protect tenant/document access across parsers, model endpoints, artifacts and review tools.
  6. Preserve exact source versions and transformation configurations so errors can be reproduced.
  7. Never publish an incomplete critical record as validated merely because one extractor succeeded.

Tip: Write the schema and acceptance rules before selecting an OCR model. “Extract the important terms” does not specify how absence, uncertainty or a multi-party agreement should be represented.

2. Capacity comes from pages and processing work

Assume a 40-page mean: 50,000 × 40 = 2M pages/month. Document arrivals average only about 0.019/s, but page work averages about 0.772/s over 30 days. Large uploads create bursts.

For an independent capacity assumption, 40% of pages require scan processing at four seconds/page:

Quantity Calculation Result
Scan pages/month 2M × 40% 800,000
Scan processing time 800,000 × 4 seconds About 889 worker-hours
Mean simultaneous scan work 800,000 / 2,592,000 × 4 About 1.24 workers
At a 10× page-arrival peak 1.24 × 10 About 12.4 workers before headroom

Ten workers processing one page every four seconds have capacity 2.5 pages/s, below the assumed 3.09 scan-pages/s peak. A queue absorbs a short burst, not sustained overload. Native parsing, extraction, model quotas and review need separate capacity models.

For a 100-page document, ten workers at four seconds/page give an idealized 40-second page-processing floor if all pages need that work. Setup, skew, rate limits and later extraction add latency. Do not divide by every worker in the fleet if the document has a lower concurrency cap.

3. Start with native text and one structured extractor

The baseline validates the file, recovers text/layout, runs one bounded structured extraction, checks the result and sends unresolved critical fields to review. Add specialized extractors or vision processing only when representative evaluation shows a benefit.

Baseline failure Improvement Benefit Cost or limitation
Native text is absent or garbled Page-level OCR/vision fallback More usable evidence Compute and recognition errors
Column/table reading order is wrong Layout-aware blocks/cells with coordinates Preserves field relationships Layout models still fail
All-fields prompt misses a difficult clause Focused extraction by field group More relevant attention/context Repeated input cost and merge conflicts
Payment value is defined in an exhibit Resolve reference graph Recovers controlling evidence Missing/cyclic references and more context
JSON is valid but value is wrong Evidence and domain validation Detects some semantic errors Source inspection may still be needed
A retry creates a duplicate downstream entry Separate source and business operation IDs Safer replay Destination-specific idempotency/reconciliation

4. Detailed architecture

Architecture / visual model
flowchart TD PDF[Authorized PDF and attachments] --> IN[Validate bytes, limits and source identity] IN --> RAW[(Immutable source versions)] RAW --> JOB[Durable page jobs with tenant and version] JOB --> TYPE{Page has usable native text?} TYPE -->|Yes| NATIVE[Extract text blocks, words and coordinates] TYPE -->|No or low quality| OCR[Evaluated OCR/layout or vision route] NATIVE --> STRUCT[Sections, tables, definitions and references] OCR --> STRUCT STRUCT --> EVID[(Versioned evidence and reference graph)] EVID --> EX[Bounded single or specialized extraction] EX --> MERGE[Field candidates with evidence] MERGE --> VALID[Schema, source and domain checks] VALID --> GATE{Critical record complete and accepted?} GATE -->|No| REVIEW[Reviewer with source page and field differences] REVIEW --> VALID GATE -->|Yes| PUB[Versioned publication transaction] PUB --> DB[(Structured records and searchable evidence)] PUB --> OUT[Optional downstream action with separate authorization]
Read diagram source
flowchart TD
    PDF[Authorized PDF and attachments] --> IN[Validate bytes, limits and source identity]
    IN --> RAW[(Immutable source versions)]
    RAW --> JOB[Durable page jobs with tenant and version]
    JOB --> TYPE{Page has usable native text?}
    TYPE -->|Yes| NATIVE[Extract text blocks, words and coordinates]
    TYPE -->|No or low quality| OCR[Evaluated OCR/layout or vision route]
    NATIVE --> STRUCT[Sections, tables, definitions and references]
    OCR --> STRUCT
    STRUCT --> EVID[(Versioned evidence and reference graph)]
    EVID --> EX[Bounded single or specialized extraction]
    EX --> MERGE[Field candidates with evidence]
    MERGE --> VALID[Schema, source and domain checks]
    VALID --> GATE{Critical record complete and accepted?}
    GATE -->|No| REVIEW[Reviewer with source page and field differences]
    REVIEW --> VALID
    GATE -->|Yes| PUB[Versioned publication transaction]
    PUB --> DB[(Structured records and searchable evidence)]
    PUB --> OUT[Optional downstream action with separate authorization]

Keep intake, page work, extraction and review independently observable. A page that fails parsing is a coverage gap; it must not disappear from the document's manifest. Publication validates the current source and lifecycle version so an obsolete job cannot overwrite a corrected record.

API and records

POST /documents
{upload_reference, request_id, document_family, attachments?} → document_id, version

GET /documents/{id}/extractions/{version}
→ processing_status, page_coverage, fields[], validation_issues[], review_status

POST /extractions/{id}/review
{expected_revision, corrections, evidence_refs, decision} → reviewed_revision
Record Essential fields
Source manifest Tenant/document/version, immutable bytes/hash, attachments and page count
Page result Physical page index, printed page label, parser/OCR version, dimensions, coordinate transform and quality/coverage
Evidence span Source version, page/cell/span reference, extracted text and source crop/coordinates
Field candidate Field ID, raw/normalized value, status, evidence references and extractor configuration
Validation result Rule/version, affected fields, outcome, evidence and severity
Review decision Reviewer, previous/new value, source evidence, decision and revision
Published record Accepted extraction revision, source lineage and downstream operation references

Use physical page index and printed page label separately: a PDF's fifth page may be labeled “iii” or “1.” A reviewer must open the correct image even after rotation or deskew.

5. Parsing, OCR and vision are choices to evaluate

PyMuPDF's text tools expose text blocks/words with position information. Plain extracted text may have an unexpected reading order. Preserve tables as cells/rows with headers and page relationships; a Markdown rendering is useful for reading but should not be the only stored representation.

Route Useful starting point Important failure cases
Native PDF extraction Reliable embedded text Hidden text, columns, repeated headers, broken encoding
OCR/layout engine Scans and repeated forms Blur, skew, handwriting, stamps, merged cells
Vision-language processing Complex visual context or difficult pages Omitted/invented text, exact digits, tiny annotations

A digital PDF can contain scanned pages or an inaccurate old OCR layer. Route and evaluate by page/content quality, not just a file-wide “native” label. Keep the original image for consequential values. If OCR reads $10,000 as $100,000, both values pass a numeric schema.

Gemini's document-understanding interface is one current multimodal option. Check the selected model's file/page/context constraints and billing. A large accepted PDF does not prove every clause was attended to accurately. Traditional OCR, dedicated document services and vision models should be compared on the actual language/layout slices, without invented universal accuracy rankings.

6. A field needs meaning, status and evidence

The following Pydantic v2 example defines a text field candidate, not the entire contract or a legal correctness test. Parties, dates, money and conditions need corresponding domain schemas; contracts may have more than two parties.

from typing import Literal
from pydantic import BaseModel, ConfigDict, Field, model_validator

class EvidenceRef(BaseModel):
    model_config = ConfigDict(extra="forbid", strict=True)
    document_version: str = Field(min_length=1)
    page: int = Field(ge=1)  # Physical, one-based page number in this API.
    span_id: str = Field(min_length=1)
    quotation: str = Field(min_length=1)

class SourcedText(BaseModel):
    model_config = ConfigDict(extra="forbid", strict=True)
    status: Literal["found", "not_found", "ambiguous", "processing_failed"]
    value: str | None
    evidence: list[EvidenceRef]

    @model_validator(mode="after")
    def consistent_candidate(self):
        if self.status == "found":
            if not self.value or not self.value.strip() or not self.evidence:
                raise ValueError("Found values require a value and evidence")
        elif self.value is not None:
            raise ValueError("Unresolved candidates must not assert a final value")
        return self

Pydantic model validators check relationships after parsing. Custom Python checks are not automatically enforced by a model provider's JSON decoder. Run them in the application after generation.

A valid evidence reference still needs validation: the document/page/span must exist, the quotation must be located in that version, and it must support the field's meaning. For scans, comparison to OCR text alone cannot detect an OCR error; inspect the image or use an independently evaluated check when required.

Field Preserve Validate
Parties Legal names, roles, aliases and source Avoid merging a parent and subsidiary without evidence
Dates Original date string, event type and normalized date if unambiguous Signing date versus effective date; ambiguous formats
Payment Decimal amount, currency, units, frequency and conditions One-off versus recurring; taxes, caps and exceptions
Obligations Actor, action, deadline, trigger and exceptions Negation, conditional duties and cross-references
Termination Rights, notice period, triggers and survival clauses Distinguish expiry, early termination and post-termination duties

An effective date after a termination date can be a warning requiring context, not a universal automatic rejection. A missing payment frequency can be valid for a one-time fee. Cross-field rules identify cases to inspect; they must not rewrite unusual contract terms to fit a template.

7. Large documents, references and parallel work

For a 200-page agreement, identify relevant sections and referenced definitions/exhibits, then fit the necessary evidence into bounded extraction contexts. Keep a manifest of included, excluded, failed and unresolved sections.

Architecture / visual model
flowchart LR DOC[Contract and authorized attachments] --> MAP[Sections and definitions] MAP --> SELECT[Candidate clauses for required fields] SELECT --> REF{Referenced controlling evidence?} REF -->|Available| FOLLOW[Read exhibit, definition or amendment] FOLLOW --> REF REF -->|Missing or cycle/budget limit| GAP[Record unresolved dependency] REF -->|Complete bounded evidence| EX[Extract field candidates] EX --> CHECK[Validate precedence, conditions and evidence] GAP --> REVIEW[Review or request missing source] CHECK --> RECORD[Record with lineage and coverage]
Read diagram source
flowchart LR
    DOC[Contract and authorized attachments] --> MAP[Sections and definitions]
    MAP --> SELECT[Candidate clauses for required fields]
    SELECT --> REF{Referenced controlling evidence?}
    REF -->|Available| FOLLOW[Read exhibit, definition or amendment]
    FOLLOW --> REF
    REF -->|Missing or cycle/budget limit| GAP[Record unresolved dependency]
    REF -->|Complete bounded evidence| EX[Extract field candidates]
    EX --> CHECK[Validate precedence, conditions and evidence]
    GAP --> REVIEW[Review or request missing source]
    CHECK --> RECORD[Record with lineage and coverage]

Track visited references and a depth/work limit. Resolve attachments through the authorized source catalog, not arbitrary links embedded in a PDF. A missing Exhibit A is an evidence gap. Do not infer its price from a template or silently exclude it to save tokens.

Amendment precedence requires the relevant agreement relationships and effective versions. A newer upload timestamp alone does not establish legal precedence. Escalate unresolved interpretation rather than making the model the final authority.

Parallel extraction changes economics and failure handling

import asyncio

async def extract_all(document, extractors):
    names = ("parties", "dates", "obligations", "termination")
    results = await asyncio.gather(*(
        extractors[name].extract_with_evidence(document) for name in names
    ), return_exceptions=True)
    return {
        name: {"status": "processing_failed"}
        if isinstance(result, BaseException)
        else {"status": "needs_validation", "fields": result}
        for name, result in zip(names, results)
    }

These are application adapters, not a claimed provider SDK. Apply per-call timeout, token reservation and concurrency limits. Retain failed categories so partial success cannot appear complete. Keep original failures in restricted diagnostics. Do not assume four extractors independently reading the same mistaken OCR will correct one another.

8. Languages, layouts and reviewer workflow

Evaluate all four languages across document families. Locale can inform interpretation, but language alone does not uniquely resolve 03/04/2026, punctuation in amounts or legal intent. Preserve the original string, units and evidence. Translation can help a reviewer, but must remain distinguishable from the signed source.

SUPPORTED_LANGUAGES = {"en", "de", "fr", "es"}

async def extract_by_language(document, language, extractors):
    if language not in SUPPORTED_LANGUAGES or language not in extractors:
        return {"status": "language_review_required", "source": document.id}
    result = await extractors[language].extract_with_evidence(document)
    return {"status": "needs_validation", "fields": result}

A registry can supply one evaluated multilingual model or specialized adapters. Handle mixed-language documents explicitly. Template/layout libraries help known forms; unfamiliar layouts should use measured fallback/review policies and become labeled regression cases after validation.

The reviewer sees each proposed field beside the exact source page, surrounding clause, linked definitions and any conflicting values. They can correct, mark absent, retain ambiguity or request missing attachments. Record the decision and its evidence; a click is not automatically a reliable training label.

Route using missing evidence, scan quality, contradictions, criticality and calibrated error estimates. Review a representative sample of unflagged fields too; an uncertain-only queue cannot measure missed errors.

9. Quality metrics and downstream safety

Precision: among returned values, how many are correct? Recall: among required source values, how many were found? Define matching and partial-credit rules before scoring. Measure exact critical amounts/dates and semantic obligations under an explicit rubric; do not hide missing clauses inside a vague “95% accuracy” score.

Measurement Why it matters
Field precision/recall by category Separates wrong values from missed values
Complete-record acceptance Shows whether the whole document is usable
Critical-field error rate Reflects high-impact mistakes
Page/reference coverage Reveals missing scans, exhibits or amendments
Review rate and minutes/record Exposes operating burden
Downstream duplicate/incorrect actions Measures consequences beyond extraction

If ten fields each have 98% correctness and errors were independent, all ten would be correct with probability 0.98^10 ≈ 81.7%. Real OCR/layout errors correlate, so measure complete-document results directly rather than multiplying marginal scores as a production estimate.

Use source IDs/content hashes for repeat processing, but a separate business identity for actions such as invoice posting. Two different scans may represent one invoice. A validated extraction must not automatically trigger an unapproved payment, and a timeout after posting needs destination reconciliation rather than blind duplication.

10. Cost and review capacity

For an illustrative native 100-page workload, assume GPT-6 Luna standard short-context rates of $0.10 input/$0.50 billed output per million tokens:

Stage Aggregate input / output tokens Cost
Section identification 50,000 / 1,000 $0.0055
Four extractors combined 100,000 / 4,000 $0.0120
Additional model validation 10,000 / 1,000 $0.0015
Text-model subtotal 160,000 / 6,000 $0.0190
Parsing/storage Hypothetical allowance $0.0300
Native partial total $0.0490

The API rate is published; the workload and infrastructure allowances are assumptions. The aggregate across four extractors is not their individual context length. Count repeated context, billed reasoning, failed calls, retries and provider-specific image/file charges.

Assuming $0.20 extra scan processing and a 60% native/40% scanned document mix, the partial machine average is $0.049 + 0.40 × $0.20 = $0.129/document. This is a costing scenario, separate from the earlier page-mix capacity example. It does not establish a fixed OCR price or prove every 200-page document stays below fifty cents.

If 10% of 50,000 documents need two minutes of review at an assumed $40/hour, review adds about $6,667/month, or $0.1333 per incoming document. Combined partial cost is about $0.2623/document, before supervision, deeper fallbacks and omitted overhead. The queue needs about 167 productive reviewer-hours/month; arrivals and skill coverage still determine staffing. Never promise 30-second review of an ambiguous contract without evidence.

Interview follow-ups

1. Why not send every PDF to one large vision model? That is a baseline to evaluate for suitable documents, but it may cost more, omit details or misread exact values. Native extraction, layout recovery and source-linked validation provide useful alternatives and diagnostics.

2. What if the price is in an exhibit? Follow the authorized reference and include the relevant source evidence. If the exhibit is missing or precedence is unclear, preserve the gap and route for review.

3. Does valid JSON mean accurate extraction? No. It establishes structure. A wrong but well-typed amount or fabricated quotation can pass schema validation, so source and domain checks remain necessary.

4. Should specialized extractors always replace one prompt? No. Compare end-to-end field/record quality, repeated input cost, latency and merge conflicts. Correlated source errors affect all extractors.

5. What does 98% field accuracy tell you about review volume? Not enough by itself. Errors across many fields and correlated document failures change complete-record acceptance. Measure the actual fraction of records requiring review and their handling time.

6. How do you handle a tenfold burst with worse scans? Bound admission and queues, separate parse/recognition failures from extraction uncertainty, reserve review capacity and prioritize by the agreed consequence policy. Do not silently lower critical-field checks to empty the queue.

7. How do you prevent duplicate accounting entries? Use business identifiers and an idempotent/transactional posting contract in addition to file deduplication. Reconcile unknown posting outcomes before retrying.

60-second interview answer

I would preserve the source and page layout, route each page through suitable native or scan processing, and extract typed candidates with explicit evidence and status. Reference resolution includes exhibits and amendments rather than assuming the first matching clause controls. Schema, source and domain checks determine what needs review, and publication retains the exact lineage. I would measure both field and complete-record quality, cost by document length, reviewer capacity and downstream action correctness before expanding document families.

Remember: Preserve → Read → Resolve references → Extract → Validate → Review and publish.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Design Support Automation That Resolves the Right Issue
NEXT LESSONDesign Movie Recommendations with Truthful Explanations →

Explore the diagram