Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Design an Image and Video Generation Platform

By Anup Rai8 min readReviewed September 2026

Interview problem: let a creative team generate, revise and export images and short videos from text and approved reference assets. Jobs can take seconds to minutes, so the product must expose progress, cancellation, version history and usage limits.

Targets below are assumptions for an interview. Rendering quality, GPU throughput and billing vary by model, resolution, duration and supported editing controls. Choose a pinned approved model after a workload-specific evaluation.

1. Requirements and scope

Functional requirements

  1. Accept a prompt, approved references, output dimensions, duration where relevant, and a generation budget.
  2. Create asynchronous jobs with visible state and optional progress events.
  3. Support variations and edits while retaining the source asset and lineage.
  4. Let users cancel work, retry eligible failures and export an approved artifact.
  5. Apply input/output policy checks and route uncertain cases for review.
  6. Record usage and enforce tenant access, retention and deletion.

Non-functional requirements

  1. Acknowledge accepted jobs within 500 ms p95, excluding large upload transfer.
  2. Target 95% completion within one minute for the agreed image class and five minutes for the agreed short-video class under admitted load.
  3. Target 99.9% monthly job-API availability; report queue and generation deadlines separately.
  4. Never publish an unapproved artifact or expose another tenant's source asset.
  5. Bound wasted GPU work after cancellation and duplicate delivery.

Defer arbitrary model training, unrestricted public hosting and guarantees of exact character consistency across every model. Ask whether users need transparent backgrounds, accurate text, temporal consistency, reproducible seeds or licensed references; each changes the model contract.

2. Capacity and storage

Assume 100,000 image jobs/day at 8 GPU-seconds each and 5,000 video jobs/day at 120 GPU-seconds each.

Quantity Calculation Implication
Image compute 800,000 GPU-seconds ≈ 222.2 GPU-hours/day Benchmark the target resolution
Video compute 600,000 GPU-seconds ≈ 166.7 GPU-hours/day Duration/steps dominate variance
Total average 388.9 / 24 ≈ 16.2 continuously busy GPUs Does not meet bursts or failover alone
At 65% target utilization 16.2 / .65 ≈ 24.9 GPUs About 25 before additional burst reserve
Images at 2 MB each 200 GB/day Retention and egress matter
Videos at 20 MB each 100 GB/day Add input references, previews and versions

Thirty days of these outputs is about 9 TB decimal before replication and deletion. Multiple candidates per job multiply both compute and output storage. A billing unit called “one generation” is not necessarily one model invocation.

3. Baseline and review

Architecture / visual model
flowchart LR C[Creator] --> API[Job API] API --> DB[(Job record)] DB --> W[Single generation worker] W --> STORE[(Private output store)] STORE --> C
Read diagram source
flowchart LR
 C[Creator] --> API[Job API]
 API --> DB[(Job record)]
 DB --> W[Single generation worker]
 W --> STORE[(Private output store)]
 STORE --> C

This is enough to validate one image model and a small workload. It is not enough for production queue recovery or safe publication.

Flaw Repair Benefit New cost
Worker crashes after writing output Immutable artifacts plus conditional job completion Recover without duplicate publication Orphan cleanup and reconciliation
Video jobs block small images Separate queues and fair scheduling Better interactive wait times Capacity balancing
A cancelled job finishes late Versioned cancellation checked before publication No unwanted asset release Some compute remains unavoidable
User edits a shared source file Content-addressed immutable input version Reproducible lineage More storage
Preview bypasses moderation One publication gate for all derivative paths Consistent release policy Preview latency

4. Detailed architecture

Architecture / visual model
flowchart TD C[Creative application] --> U[Authenticated upload and job API] U --> IN[(Private immutable input assets)] U --> P[Input validation and budget reservation] P --> J[(Job state and transactional outbox)] J --> Q[Image and video queues] Q --> S[Lease and fair scheduler] S --> I[Image worker pool] S --> V[Video worker pool] IN --> I IN --> V I --> OUT[(Unpublished candidate artifacts)] V --> OUT OUT --> CHECK[Output checks and review] CHECK --> PUB[Conditional publication] J --> PUB PUB --> CAT[(Versioned asset catalog)] CAT --> DL[Authorized short-lived download] DL --> C I --> USAGE[Usage settlement and reconciliation] V --> USAGE J --> EVENTS[Progress event stream] EVENTS --> C
Read diagram source
flowchart TD
 C[Creative application] --> U[Authenticated upload and job API]
 U --> IN[(Private immutable input assets)]
 U --> P[Input validation and budget reservation]
 P --> J[(Job state and transactional outbox)]
 J --> Q[Image and video queues]
 Q --> S[Lease and fair scheduler]
 S --> I[Image worker pool]
 S --> V[Video worker pool]
 IN --> I
 IN --> V
 I --> OUT[(Unpublished candidate artifacts)]
 V --> OUT
 OUT --> CHECK[Output checks and review]
 CHECK --> PUB[Conditional publication]
 J --> PUB
 PUB --> CAT[(Versioned asset catalog)]
 CAT --> DL[Authorized short-lived download]
 DL --> C
 I --> USAGE[Usage settlement and reconciliation]
 V --> USAGE
 J --> EVENTS[Progress event stream]
 EVENTS --> C

A model result is a candidate artifact. The publication service checks the current job state, access and review outcome before making it available.

5. Data and API contracts

POST /jobs accepts idempotency_key, input_asset_versions, prompt, model_profile, seed where supported, output settings and a candidate limit. Return 202 with a job ID and status URL. POST /jobs/{id}/cancel changes intent; it cannot promise to undo already charged computation. Downloads authorize the actual asset version and tenant.

Record Fields Invariant
Job tenant, ID, state, version, deadline, reservation, input versions Valid state transitions use compare-and-swap
Attempt job ID, lease token, model/runtime revision, settings, output reference A stale worker cannot publish
Asset content hash, private object key, media metadata, source lineage, review state Immutable bytes; access checked at delivery
Publication job version, selected asset, approving actor/checks Cancelled or superseded work cannot become current

Valid states include accepted, queued, running, checking, ready, failed and cancelled. Use a fencing token on leases; an expired worker may still be running, so a lease timeout alone does not guarantee exclusive publication.

6. Trace the generation and edit paths

  1. Authenticate, validate reference access and pin immutable input versions.
  2. Validate media type, dimensions, duration, policy and maximum cost.
  3. Commit the job and outbox event together; a dispatcher publishes work to the queue.
  4. A worker claims a lease, checks cancellation and runs the pinned model profile.
  5. Write candidates under immutable attempt-specific keys; report measured work.
  6. Validate output and confirm the current job version before publication.
  7. Deliver through authorized download URLs with bounded lifetime; apply cache and revocation rules.
  8. An edit creates a new job referencing the earlier asset version. Preserve both until retention or user deletion removes them.

A fixed seed can help reproduce an output with the same supported configuration; it is not a universal bit-for-bit guarantee across hardware, runtimes or provider upgrades.

7. Failure and quality evaluation

Test Expected behavior
Duplicate queue delivery Claim/version checks prevent a second current publication
Worker completes after cancellation Candidate remains unpublished and is cleaned up
Safety checker unavailable No public release until required checks succeed
Reference access revoked Revalidate before use/release under the agreed revocation policy
Corrupt video container Validation fails the attempt; no broken download advertised
Browser disconnects Job continues under its contract; client reconnects by job ID

Evaluate prompt adherence, editing preservation, text legibility where required, temporal stability, policy false positives/negatives, and human acceptance. Report by media class. One overall “quality score” can hide unusable video or a broken editing path.

8. Cost-benefit and closing

At an illustrative $2/GPU-hour, 388.9 busy GPU-hours/day is about $777.8/day of utilized compute. A 25-GPU reserved fleet costs 25 × 24 × $2 = $1,200/day before burst reserve. Add storage, downloads, checks and operator effort. Compare pay-per-job providers on the same workload and approved data-processing terms.

Choice Benefit Tradeoff
Draft at low resolution Cheap composition feedback Final render may differ
More candidates Better chance of a usable result Multiplies cost; selection still needed
Dedicated video pool Predictable image latency Idle capacity
Spot/preemptible workers Lower eligible batch compute cost Lost work and deadline risk

Q1: Does exactly-once queue delivery solve duplicate generation?

Sample answer: No. A worker can crash after a side effect but before recording completion. Use idempotent job creation, durable attempts, immutable artifact keys and fenced conditional publication. Some duplicate compute may remain, but duplicate release and charging can be reconciled.

Q2: What should cancellation guarantee?

Sample answer: Stop queued work, request cancellation of running work, and prevent later publication once cancellation wins the state transition. Define how already consumed work is charged; do not promise to reverse GPU time.

Q3: How do you close the design?

Sample answer: Treat generation as an asynchronous, versioned artifact workflow. Separate compute completion from publication, protect input/output access and measure accepted creative outcomes. Add model variety only when the asset and job lifecycle is reliable.

Recall: Pin inputs → Reserve budget → Lease work → Validate candidates → Publish conditionally → Retain lineage.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Design a Deadline-Aware Batch Inference Platform
NEXT LESSONDesign a Multi-Tenant Model-Serving Platform →

Explore the diagram