Interview problem: let a creative team generate, revise and export images and short videos from text and approved reference assets. Jobs can take seconds to minutes, so the product must expose progress, cancellation, version history and usage limits.
Targets below are assumptions for an interview. Rendering quality, GPU throughput and billing vary by model, resolution, duration and supported editing controls. Choose a pinned approved model after a workload-specific evaluation.
1. Requirements and scope
Functional requirements
- Accept a prompt, approved references, output dimensions, duration where relevant, and a generation budget.
- Create asynchronous jobs with visible state and optional progress events.
- Support variations and edits while retaining the source asset and lineage.
- Let users cancel work, retry eligible failures and export an approved artifact.
- Apply input/output policy checks and route uncertain cases for review.
- Record usage and enforce tenant access, retention and deletion.
Non-functional requirements
- Acknowledge accepted jobs within 500 ms p95, excluding large upload transfer.
- Target 95% completion within one minute for the agreed image class and five minutes for the agreed short-video class under admitted load.
- Target 99.9% monthly job-API availability; report queue and generation deadlines separately.
- Never publish an unapproved artifact or expose another tenant's source asset.
- Bound wasted GPU work after cancellation and duplicate delivery.
Defer arbitrary model training, unrestricted public hosting and guarantees of exact character consistency across every model. Ask whether users need transparent backgrounds, accurate text, temporal consistency, reproducible seeds or licensed references; each changes the model contract.
2. Capacity and storage
Assume 100,000 image jobs/day at 8 GPU-seconds each and 5,000 video jobs/day at 120 GPU-seconds each.
| Quantity | Calculation | Implication |
|---|---|---|
| Image compute | 800,000 GPU-seconds ≈ 222.2 GPU-hours/day | Benchmark the target resolution |
| Video compute | 600,000 GPU-seconds ≈ 166.7 GPU-hours/day | Duration/steps dominate variance |
| Total average | 388.9 / 24 ≈ 16.2 continuously busy GPUs | Does not meet bursts or failover alone |
| At 65% target utilization | 16.2 / .65 ≈ 24.9 GPUs | About 25 before additional burst reserve |
| Images at 2 MB each | 200 GB/day | Retention and egress matter |
| Videos at 20 MB each | 100 GB/day | Add input references, previews and versions |
Thirty days of these outputs is about 9 TB decimal before replication and deletion. Multiple candidates per job multiply both compute and output storage. A billing unit called “one generation” is not necessarily one model invocation.
3. Baseline and review
Read diagram source
flowchart LR
C[Creator] --> API[Job API]
API --> DB[(Job record)]
DB --> W[Single generation worker]
W --> STORE[(Private output store)]
STORE --> C
This is enough to validate one image model and a small workload. It is not enough for production queue recovery or safe publication.
| Flaw | Repair | Benefit | New cost |
|---|---|---|---|
| Worker crashes after writing output | Immutable artifacts plus conditional job completion | Recover without duplicate publication | Orphan cleanup and reconciliation |
| Video jobs block small images | Separate queues and fair scheduling | Better interactive wait times | Capacity balancing |
| A cancelled job finishes late | Versioned cancellation checked before publication | No unwanted asset release | Some compute remains unavoidable |
| User edits a shared source file | Content-addressed immutable input version | Reproducible lineage | More storage |
| Preview bypasses moderation | One publication gate for all derivative paths | Consistent release policy | Preview latency |
4. Detailed architecture
Read diagram source
flowchart TD
C[Creative application] --> U[Authenticated upload and job API]
U --> IN[(Private immutable input assets)]
U --> P[Input validation and budget reservation]
P --> J[(Job state and transactional outbox)]
J --> Q[Image and video queues]
Q --> S[Lease and fair scheduler]
S --> I[Image worker pool]
S --> V[Video worker pool]
IN --> I
IN --> V
I --> OUT[(Unpublished candidate artifacts)]
V --> OUT
OUT --> CHECK[Output checks and review]
CHECK --> PUB[Conditional publication]
J --> PUB
PUB --> CAT[(Versioned asset catalog)]
CAT --> DL[Authorized short-lived download]
DL --> C
I --> USAGE[Usage settlement and reconciliation]
V --> USAGE
J --> EVENTS[Progress event stream]
EVENTS --> C
A model result is a candidate artifact. The publication service checks the current job state, access and review outcome before making it available.
5. Data and API contracts
POST /jobs accepts idempotency_key, input_asset_versions, prompt, model_profile, seed where supported, output settings and a candidate limit. Return 202 with a job ID and status URL. POST /jobs/{id}/cancel changes intent; it cannot promise to undo already charged computation. Downloads authorize the actual asset version and tenant.
| Record | Fields | Invariant |
|---|---|---|
| Job | tenant, ID, state, version, deadline, reservation, input versions | Valid state transitions use compare-and-swap |
| Attempt | job ID, lease token, model/runtime revision, settings, output reference | A stale worker cannot publish |
| Asset | content hash, private object key, media metadata, source lineage, review state | Immutable bytes; access checked at delivery |
| Publication | job version, selected asset, approving actor/checks | Cancelled or superseded work cannot become current |
Valid states include accepted, queued, running, checking, ready, failed and cancelled. Use a fencing token on leases; an expired worker may still be running, so a lease timeout alone does not guarantee exclusive publication.
6. Trace the generation and edit paths
- Authenticate, validate reference access and pin immutable input versions.
- Validate media type, dimensions, duration, policy and maximum cost.
- Commit the job and outbox event together; a dispatcher publishes work to the queue.
- A worker claims a lease, checks cancellation and runs the pinned model profile.
- Write candidates under immutable attempt-specific keys; report measured work.
- Validate output and confirm the current job version before publication.
- Deliver through authorized download URLs with bounded lifetime; apply cache and revocation rules.
- An edit creates a new job referencing the earlier asset version. Preserve both until retention or user deletion removes them.
A fixed seed can help reproduce an output with the same supported configuration; it is not a universal bit-for-bit guarantee across hardware, runtimes or provider upgrades.
7. Failure and quality evaluation
| Test | Expected behavior |
|---|---|
| Duplicate queue delivery | Claim/version checks prevent a second current publication |
| Worker completes after cancellation | Candidate remains unpublished and is cleaned up |
| Safety checker unavailable | No public release until required checks succeed |
| Reference access revoked | Revalidate before use/release under the agreed revocation policy |
| Corrupt video container | Validation fails the attempt; no broken download advertised |
| Browser disconnects | Job continues under its contract; client reconnects by job ID |
Evaluate prompt adherence, editing preservation, text legibility where required, temporal stability, policy false positives/negatives, and human acceptance. Report by media class. One overall “quality score” can hide unusable video or a broken editing path.
8. Cost-benefit and closing
At an illustrative $2/GPU-hour, 388.9 busy GPU-hours/day is about $777.8/day of utilized compute. A 25-GPU reserved fleet costs 25 × 24 × $2 = $1,200/day before burst reserve. Add storage, downloads, checks and operator effort. Compare pay-per-job providers on the same workload and approved data-processing terms.
| Choice | Benefit | Tradeoff |
|---|---|---|
| Draft at low resolution | Cheap composition feedback | Final render may differ |
| More candidates | Better chance of a usable result | Multiplies cost; selection still needed |
| Dedicated video pool | Predictable image latency | Idle capacity |
| Spot/preemptible workers | Lower eligible batch compute cost | Lost work and deadline risk |
Q1: Does exactly-once queue delivery solve duplicate generation?
Sample answer: No. A worker can crash after a side effect but before recording completion. Use idempotent job creation, durable attempts, immutable artifact keys and fenced conditional publication. Some duplicate compute may remain, but duplicate release and charging can be reconciled.
Q2: What should cancellation guarantee?
Sample answer: Stop queued work, request cancellation of running work, and prevent later publication once cancellation wins the state transition. Define how already consumed work is charged; do not promise to reverse GPU time.
Q3: How do you close the design?
Sample answer: Treat generation as an asynchronous, versioned artifact workflow. Separate compute completion from publication, protect input/output access and measure accepted creative outcomes. Add model variety only when the asset and job lifecycle is reliable.
Recall: Pin inputs → Reserve budget → Lease work → Validate candidates → Publish conditionally → Retain lineage.