Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Design a Multi-Tenant Contract-Analysis Platform

By Anup Rai13 min readReviewed September 2026

Interview problem: let competing businesses upload private contracts and ask grounded questions on shared infrastructure, while enforcing access, predictable service and accountable data lifecycle operations.

A tenant is a customer workspace with its own members, data and policy. Multi-tenancy lets multiple tenants share parts of the service. Data isolation prevents unauthorized access; resource isolation prevents one customer's workload from consuming another's promised capacity. Neither follows automatically from adding a tenant_id column.

This is a hypothetical Learnastra interview scenario. Quantities and budgets are assumptions. The interview should explain enforceable boundaries and evidence, rather than promising that any architecture has literally zero risk.

1. Requirements and boundaries

Clarify whether users can belong to multiple tenants, whether contracts have narrower team permissions, which regions/providers are permitted, and which customers need dedicated resources. Tenant administration, document access and billing administration can be different roles.

Functional requirements

  1. Authenticate members and authorize the active workspace from trusted membership records.
  2. Upload, parse, version and index contracts with tenant and document permissions.
  3. Answer questions using only evidence that the caller is allowed to access.
  4. Support citations, conversation history and generated artifacts under the same access rules.
  5. Offer administrator-controlled exports, deletion/offboarding and status reports.
  6. Meter usage and enforce tenant service tiers and budgets.
  7. Support migration between shared and dedicated capacity without mixing customer data.

Nonfunctional requirements

  1. Support 500 customers with 10,000–100,000 documents each at initial scale.
  2. Deny unauthorized cross-tenant and within-tenant document access on every path.
  3. Target p95 complete checked answers below two seconds for a bounded routine query class; complex analysis has a separate deadline.
  4. Propose 99.9% monthly query availability and define behavior during authorization or provider outages.
  5. Keep tenant-specific workloads within concurrency, token, storage and background-job allowances.
  6. Encrypt data in transit/at rest, protect keys, restrict telemetry and honor the approved retention/residency policy.
  7. Produce evidence for agreed control assessments and applicable data-rights processes; selecting a cloud product does not establish compliance by itself.

Tip: Ask “which operation, by which identity, against which object?” A tenant ID in a valid session still does not authorize every document belonging to that tenant.

2. Size documents, vectors and query load separately

For planning, assume an average 50,000 documents/customer, eight passages/document and 10KB normalized text/document.

Quantity Calculation Result
Documents 500 × 50,000 25M
Normalized text 25M × 10KB 250GB
Passages 25M × 8 200M
Raw 1,024-dimensional float32 vectors 200M × 1,024 × 4 bytes 819.2GB
Two total vector copies 819.2GB × 2 1.6384TB

These decimal-unit figures exclude original PDFs, index graphs, payloads, versions and backups. Long contracts and tables may produce far more passages. A vector-service budget must be checked against the actual filter/index workload and replication, not just the number of tenants.

At 10M routine queries/month over 30 days, average load is about 3.86/s. A 10× peak is about 38.6/s. At 1.2 seconds mean residence, that implies about 46 in-flight queries under stable assumptions. Tenant traffic is usually uneven; averages must not hide a dominant customer or a large ingestion backfill.

3. Start with a shared, explicitly scoped service

The baseline uses an authenticated API, a relational catalog, object storage, one search store and a bounded generation service. Every data access receives trusted tenant/user scope. Start with a shared deployment when its isolation and capacity controls satisfy the actual contract.

Baseline failure Added control Benefit Cost or remaining limit
An application query omits tenant filtering Independently enforced database policies or scoped service boundary Contains some query-building mistakes Must cover each store and actual credentials
Cached answer crosses users/tenants Scope plus evidence/access version in cache identity Prevents unsafe reuse Lower hit rate and revocation handling
Background worker trusts a forged tenant label Trusted job envelope plus current authorization/lifecycle checks Protects asynchronous paths Credential and retry-state design
One backfill fills the shared queue Admission limits and separate scheduling classes Protects interactive requests Reserved/idle capacity
A tenant needs stronger separation Dedicated deployment with scoped credentials Smaller resource/blast-radius boundary Provisioning, upgrades and cost
A delayed job recreates deleted data Lifecycle fence, drain and reconciliation Prevents resurrection Cross-store coordination

A prompt that says “answer only for this tenant” can guide phrasing, but cannot enforce a data boundary. By the time unauthorized text reaches the model, confidentiality has already failed.

4. Detailed query architecture

Architecture / visual model
flowchart TD U[User and selected workspace] --> ID[Verify identity and current membership] ID --> POLICY[Resolve document rights and lifecycle policy] POLICY --> ADMIT[Tenant quota and bounded admission] ADMIT --> ROUTE[Trusted deployment directory] ROUTE -->|Shared| SG[Scoped search gateway] ROUTE -->|Dedicated| DG[Tenant deployment and scoped credentials] SG --> SS[(Shared indexes)] DG --> DS[(Dedicated indexes)] SS --> CAND[Candidate IDs and evidence revisions] DS --> CAND CAND --> AUTH[Authoritative document-access validation] CAT[(Tenant, document and policy catalog)] --> AUTH AUTH --> READ[Read only permitted source passages] READ --> GEN[Approved model endpoint and bounded generation] GEN --> VALIDATE[Evidence, current rights and output checks] VALIDATE --> RESULT[Answer, citations and protected artifacts] VALIDATE --> AUDIT[Restricted audit metadata and usage ledger]
Read diagram source
flowchart TD
    U[User and selected workspace] --> ID[Verify identity and current membership]
    ID --> POLICY[Resolve document rights and lifecycle policy]
    POLICY --> ADMIT[Tenant quota and bounded admission]
    ADMIT --> ROUTE[Trusted deployment directory]
    ROUTE -->|Shared| SG[Scoped search gateway]
    ROUTE -->|Dedicated| DG[Tenant deployment and scoped credentials]
    SG --> SS[(Shared indexes)]
    DG --> DS[(Dedicated indexes)]
    SS --> CAND[Candidate IDs and evidence revisions]
    DS --> CAND
    CAND --> AUTH[Authoritative document-access validation]
    CAT[(Tenant, document and policy catalog)] --> AUTH
    AUTH --> READ[Read only permitted source passages]
    READ --> GEN[Approved model endpoint and bounded generation]
    GEN --> VALIDATE[Evidence, current rights and output checks]
    VALIDATE --> RESULT[Answer, citations and protected artifacts]
    VALIDATE --> AUDIT[Restricted audit metadata and usage ledger]

The deployment directory maps a trusted tenant to its resources and version. A user cannot choose an arbitrary collection, bucket prefix, region or model credential. Dedicated resources still need permissions among users within the customer.

API and records

POST /workspaces/{workspace_id}/documents
→ document_id, version, ingestion_status

POST /workspaces/{workspace_id}/answers
{query, allowed_document_selection?} → answer, citations, coverage

POST /workspaces/{workspace_id}/exports
{scope, approved_destination_id} → export_job_id

POST /workspaces/{workspace_id}/deletion-requests
{scope, request_id} → deletion_job_id

The route names the requested workspace; current membership and object authorization determine whether it may be used. A client-provided document selection can narrow access, never expand it.

Record Essential fields
Membership User, tenant, role, status and authorization version
Tenant deployment Tenant, region, service tier, resource references and routing generation
Document revision Tenant, document, version, ACL reference, lifecycle state and object reference
Job Trusted tenant/object scope, initiating actor, policy version, deadline and lifecycle fence
Usage reservation Tenant, request/attempt, reserved and settled token/compute quantities
Artifact/export Tenant, permitted audience, source versions, expiry and destination
Deletion task Approved scope, affected stores, per-store proof/status and retention exceptions

Use tenant-qualified IDs and constraints where appropriate. Avoid error messages, global uniqueness checks or search counts that reveal the existence of another customer's records.

5. Choose isolation by its actual boundary

Arrangement What it separates What it still shares When it may fit
Shared collection with tenant payload filters Logical search scope through enforced access service Index process, compute and broad service credential Smaller workloads with accepted shared controls
Separate collection and collection-scoped credentials Collection access and configuration Deployment resources and operational plane Different index configuration or credential boundary
Dedicated deployment Service process/capacity and scoped data credentials Possibly cloud account, control plane and upstream provider quotas Contractual, residency, workload or recovery needs

Document count alone is not an isolation policy. Twenty customers may need dedicated capacity because of traffic or contracts even if none has more than 100,000 documents. Define measured migration triggers and the customer's agreed boundary.

Qdrant's tenant payload index option, is_tenant, helps organize data for tenant-local access patterns; it is not an authorization switch. Its granular access tokens support collection-level permissions. A token allowed to read a shared collection must not be handed to a tenant as though it restricted access to that tenant's payload rows.

Keep shared-store credentials in a trusted backend, forbid unscoped APIs and cover search, point lookup, scroll, count, snapshots and administrative operations. Candidate metadata validation can prevent unauthorized text delivery even if a query filter fails, but it does not excuse a store exposure or leaked IDs.

PostgreSQL row-level security, with its limits

Row-level security (RLS) restricts rows visible or writable by a database role according to policies. Enable it on relevant tables and test both read policies and write checks. PostgreSQL uses default deny when RLS is enabled without an applicable policy. Superusers and BYPASSRLS roles bypass it; table owners normally do too unless FORCE ROW LEVEL SECURITY applies. PostgreSQL policy documentation.

Use a non-owner application role and a trusted, transaction-scoped tenant context. Connection pools must reset context reliably. If arbitrary client SQL can set the tenant context, a custom session variable is not an independent security boundary. Protect privileged helpers and audit policy combinations; a permissive extra policy can broaden access.

RLS in PostgreSQL does not protect an independent vector store, object store or cache. Tenant scope also does not replace document ACLs. Validate revocations at the defined delivery boundary and document the consistency contract for in-flight requests.

6. Ingestion: identity and lifecycle follow the data

Architecture / visual model
flowchart LR UP[Authenticated scoped upload] --> LIMIT[Type, size, rights and lifecycle checks] LIMIT --> OBJ[(Quarantined tenant-scoped source object)] OBJ --> JOB[Durable trusted job with document version] JOB --> PARSE[Isolated parsing and extraction] PARSE --> CHUNK[Versioned passages and ACL references] CHUNK --> EMB[Approved tenant/region embedding endpoint] EMB --> INDEX[(Scoped vector and lexical entries)] CHUNK --> TEXT[(Protected passage text)] INDEX --> READY[Visibility checks and current lifecycle fence] TEXT --> READY READY --> PUB[Publish searchable document revision]
Read diagram source
flowchart LR
    UP[Authenticated scoped upload] --> LIMIT[Type, size, rights and lifecycle checks]
    LIMIT --> OBJ[(Quarantined tenant-scoped source object)]
    OBJ --> JOB[Durable trusted job with document version]
    JOB --> PARSE[Isolated parsing and extraction]
    PARSE --> CHUNK[Versioned passages and ACL references]
    CHUNK --> EMB[Approved tenant/region embedding endpoint]
    EMB --> INDEX[(Scoped vector and lexical entries)]
    CHUNK --> TEXT[(Protected passage text)]
    INDEX --> READY[Visibility checks and current lifecycle fence]
    TEXT --> READY
    READY --> PUB[Publish searchable document revision]

The server derives scope at upload and persists it in authenticated job records; workers revalidate it at sensitive steps. A tenant or document can be suspended/deleted while a job is queued. The worker's captured tenant ID alone does not authorize a later write.

  1. Validate content and quotas before expensive parsing, then quarantine untrusted files.
  2. Bind source objects, chunks, embeddings and logs to a stable document revision.
  3. Use an approved processing endpoint for every data-bearing stage, including embeddings and OCR.
  4. Publish only after required writes are searchable and the document is still eligible.
  5. Fence writes/publication against the current lifecycle generation; drain or reject stale workers.
  6. Reconcile incomplete writes and clean orphaned artifacts.

Encrypt storage and transport while retaining the fields needed for authorized filtering. Opaque application encryption of a field prevents ordinary equality search unless the system deliberately implements a suitable searchable representation and accepts its leakage tradeoffs. Embeddings can reveal source information and require data protection too.

7. Caches, sessions, logs and provider boundaries

A safe answer-cache identity includes tenant, permitted audience or user scope, query, source revisions, authorization version and model/prompt configuration. Validate current rights on retrieval. An answer created for a tenant administrator cannot automatically be reused for a restricted member.

Keep conversation state and exports under object authorization, not just hard-to-guess URLs. For object delivery, use an authenticated endpoint or a narrowly scoped, expiring mechanism consistent with revocation requirements. Never mark private artifacts publicly cacheable.

Provider retention/training terms, region and subprocessors are part of the tenant contract. A private deployment or tenant-specific adapter can satisfy particular requirements, but fine-tuning is not a privacy proof and may memorize data. Do not fine-tune a pooled model on customer contracts without the appropriate authorization and design.

Audit metadata should identify actor, tenant, object, action, decision and version. Avoid copying complete contracts and prompts into ordinary billing logs. Sensitive investigations may need separately authorized evidence retention and access.

8. Prevent noisy neighbors and explain the two-second target

Suppose tenant A starts a large backfill while tenant B requests short answers. One FIFO queue makes B wait behind A; autoscaling workers may only exhaust a shared provider quota faster.

  1. Separate interactive, ingestion, export and evaluation work classes.
  2. Enforce per-tenant concurrent work, queued bytes/items, token reservations and storage limits.
  3. Use fair scheduling and reserved capacity appropriate to service commitments.
  4. Charge retries and failed attempts to the originating tenant's resource budget.
  5. Reject or defer excess work clearly; bound queue age rather than accepting an unlimited backlog.
  6. Monitor completion and oldest queued work by tenant and class, not just fleet averages.

An illustrative two-second answer budget is 100ms identity/admission, 300ms retrieval, 200ms reranking, 1,100ms short generation, 200ms checks and 100ms headroom. This is a target to test, not a model speed claim. Long contracts or deeper comparisons should enter a different bounded query class. Measure the whole request; component percentiles do not add into an end-to-end percentile.

The operational lesson is identity follows data; budget follows work. See handling overload and access-control design.

9. Deletion, export and control evidence

SOC 2 Type II examines relevant controls and their operating effectiveness over a period; it is not a checklist certificate or a claim that every system is secure. The scope, evidence and exceptions matter. AICPA SOC 2 resources.

Control objective Engineering evidence
Access is authorized Membership lifecycle, role configuration, access reviews and negative tests
Data is protected Encryption/key policy, endpoint configuration and credential rotation evidence
Changes are controlled Reviewed deployments, test results and rollback records
Service recovers Restore drills, measured recovery objectives and incident follow-up
Data lifecycle is accountable Retention inventory, deletion/export jobs and verified outcomes

These are design examples, not a complete audit program. A TLS version or quarterly report alone cannot establish control effectiveness.

Do not confuse tenant offboarding with an individual's erasure request

An offboarding request may remove a customer's workspace; an individual's request may cover only eligible personal data within many records. GDPR Article 17 provides erasure rights with grounds and exceptions. The responsible legal/data owner defines the approved scope and applicable retention. An internal “deletion certificate” cannot override those requirements or prove bytes disappeared from every copy.

  1. Authenticate the requester, resolve scope/authority and record applicable exceptions.
  2. Mark affected objects unavailable for ordinary serving and issue a lifecycle fence against new derived writes.
  3. Drain or invalidate in-flight jobs, then erase eligible source objects, passages, vectors, sessions, caches and exports.
  4. Apply approved retention/restriction to remaining audit or held records; do not assume all logs must remain forever.
  5. Track backups and subprocessors explicitly; retain an erasure ledger so restoring a backup reapplies deletions before serving.
  6. Verify each store, retry failed work and report pending copies or exceptions accurately.

A backup-expiry request is not confirmed erasure. The retention plan must specify deadlines, restricted use and restore behavior; operational inconvenience is not a blanket legal exemption.

Executable example: honest completion reporting

def erasure_status(required_stores, states):
    if not required_stores:
        raise ValueError("Approved plan must identify affected stores")
    terminal = {"erased", "restricted_exception"}
    pending = sorted(store for store in required_stores
                     if states.get(store) not in terminal)
    if pending:
        return {"status": "pending", "stores": pending}
    retained = sorted(store for store in required_stores
                      if states[store] == "restricted_exception")
    return {"status": "completed_with_retained_records" if retained else "erased",
            "retained_stores": retained}

The inputs come from an authorized deletion plan and verified worker results. restricted_exception requires a documented, approved retention basis and actual restriction; a worker cannot assign it merely because deletion failed. Missing or scheduled-backup results stay pending.

Export and dedicated-capacity migration

Authorize the export's scope, destination and audience, then create a protected manifest with checksums, expiry and completion status. The product may export agreed documents and metadata without exporting internal embeddings or model artifacts. GDPR portability has its own scope and conditions; it is not automatically a right to every internal derived representation.

For a tenant migration, copy a consistent snapshot, replay changes, validate tenant/document counts and content hashes, then switch a versioned routing entry. Fence old writers and reconcile in-flight work. Remove obsolete source copies under policy after validation; rollback must not restore revoked access or resurrect deleted records.

10. Economics and release checks

At 10M routine calls/month, 2,000 input and 300 billed output tokens/call, GPT-6 Luna standard short-context rates of $0.10/$0.50 per million give $3,500/month in generation. This is a current candidate to evaluate, not a promise that complex legal analysis fits the fast model or latency target. API pricing.

Component Illustrative monthly allowance
Shared vector/search capacity $2,500
Dedicated deployments for 20 tenants $4,000
Routine generation $3,500
Object storage $1,500
Restricted audit logging $500
Partial total $12,000
Mean across 500 tenants $24

Infrastructure allowances require measured sizing and are not cloud quotes. Include ingestion, embeddings, backups, deeper-model calls, access services, support and control operations. The mean is not the cost of each customer. Allocate shared fixed capacity plus attributable usage and support; avoid hiding a loss-making large tenant behind the average. At a hypothetical $100 per independent deployment, 500 deployments would cost $50,000 before other services, but the comparison only matters if both options meet the same requirements.

Before rollout, test:

  1. A deliberate Customer A request for Customer B data through search, direct IDs, cache, export, jobs and citations.
  2. Same-tenant users with different document permissions and a revocation during generation.
  3. Missing tenant context, connection-pool reuse and actual RLS bypass/owner roles.
  4. One tenant exhausting queues/tokens while another uses contracted capacity.
  5. Offboarding racing with indexing, export and provider retries.
  6. Restore from a backup containing erased records and apply the erasure ledger before serving.
  7. Shared-to-dedicated migration with traffic, deletion and permission changes in progress.

Security owns the tested boundary, platform operations owns capacity/recovery, and data/legal owners define retention and export obligations. Keep an incident process for suspected leakage: inspect actual retrieval, cache, conversation and artifact provenance, contain access if warranted, and preserve scoped evidence.

Interview follow-ups

1. What if an ORM omits the tenant filter? An independently enforced database policy or access service should reject unauthorized access. Test the real role and each store. A second optional filter in the same faulty code path is weak defense in depth.

2. Does a dedicated collection solve isolation? It can provide a collection-level credential boundary and separate configuration. It may still share compute and control-plane access, and it does not establish document permissions within the tenant.

3. Can a tenant-specific prompt prevent leakage? No. Access must be enforced before evidence reaches the model and before artifacts are delivered. Prompts cannot protect data already supplied to the wrong context.

4. What if generated text resembles a competitor's confidential contract? Investigate provenance and actual transmissions. Matching text alone does not establish its source, but it also does not justify dismissing a potential incident. Check every data path and contain exposure when warranted.

5. When should a tenant move to dedicated resources? When measured workload, isolation, residency, recovery or contractual requirements justify the extra cost. Use a verified migration and routing process rather than an arbitrary document threshold.

6. When can deletion be reported complete? When the approved scope has verified outcomes across affected stores, with retained exceptions disclosed and pending copies identified. Enqueuing backup expiry or deleting the vector index is insufficient.

60-second interview answer

I would derive tenant and document authority from trusted membership and policy records, then enforce it across retrieval, storage, jobs, caches and delivery. Shared infrastructure needs independent access controls and fair admission; dedicated resources address specific boundaries rather than removing all risk. Versioned ingestion and lifecycle fences prevent stale jobs from recreating deleted data. I would verify export, migration and restore behavior, and measure tenant-level latency, cost and interference alongside evidence for the agreed control obligations.

Remember: Identity → Object rights → Fair capacity → Protected delivery → Verified lifecycle.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Design an Agent That Produces a Reviewable Code Change
NEXT LESSONDesign Support Automation That Resolves the Right Issue →

Explore the diagram