Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Multimodal Generation: From Prompt to Publishable Media

By Anup Rai27 min readReviewed September 2026

Multimodal generation produces content in one or more media types—such as images, video or audio—conditioned on text or other media. A text-to-image model is one example. A system that generates a video with dialogue from a script and reference pictures is another. Multimodal understanding interprets existing media; generation creates or edits media. A product can do both.

In an interview, the central question is: How does a requested asset become a usable, authorized and affordable result, even when generation is slow or fails? A successful model call is only one step.

This chapter develops the concepts first, then designs a 30-second video service from requirements through recovery, evaluation and cost. Workload and cost figures in the design are explicit interview assumptions. Product availability was checked on September 24, 2026.

Learn the vocabulary first

Term Standard meaning Concrete example
Modality A type of information or representation Text, image, audio and video
Conditioning Input information that guides a generation A product photograph and a description of the camera movement
Text-to-image / image-to-video Generation whose name describes its input and output Animate a supplied reference image into a short clip
Joint generation A model produces related modalities together Dialogue and video generated as one coordinated output
Cascade Separate stages produce or transform successive outputs Generate narration, generate video, then align and compose them
Latent representation A learned internal encoding, often more compact than raw media An image encoded before iterative generation in latent space
Seed An initial value for a pseudorandom process Hold a randomness setting constant while comparing prompts
Inference steps Iterations used by an iterative generation procedure A draft uses fewer denoising or integration steps, if supported
Keyframe A specified frame used to guide part of a video Require the shot to begin on the product photograph
Rendition A particular encoded or resized version of an asset A vertical MP4 and a square MP4 derived from one approved edit
Muxing Combining encoded media streams in a container Put video and audio streams into an MP4 without necessarily re-encoding them
Provenance Information about an asset's origin and changes Which source image, model and editing operation produced a rendition

Memory card: conditioning defines the requested result; generation proposes pixels or samples; validation determines whether the result can be used.

How generation works at a useful interview depth

  1. Encode the inputs. Text, images or audio become representations the model can process. Inputs still need file, size, authorization and content checks.
  2. Generate a representation. A diffusion model learns to reverse a noising process. A flow-based model learns a vector field that transports a simple distribution toward a data distribution. An autoregressive model predicts the next element conditional on preceding elements. These approaches can appear in hybrid systems.
  3. Decode and post-process. Convert the result to pixels, waveform samples or another usable representation; resize, encode or compose when required.
  4. Check the actual output. A valid prompt does not guarantee legible text, correct product geometry, safe imagery, an accurate voice or usable timing.

“Transformer” and “diffusion” do not describe mutually exclusive model categories: a transformer can be the neural network used inside a diffusion or flow-based generator. Similarly, native audio-video support does not remove the need for editing, evaluation or rights checks. The original latent diffusion paper and flow matching paper explain the training distinctions.

How to control an image or shot

Technique What it controls What it does not guarantee
Text and, where supported, negative prompts Desired content and discouraged attributes Exact compliance or supported negative-prompt semantics on every API
ControlNet-style conditioning Structural signals such as pose, edges or depth Identity, ownership or compatibility with every base model
Reference-image adapter Visual guidance from example images Exact preservation of a face, brand mark or product
Inpainting Regenerates a selected region using surrounding context Perfectly unchanged pixels outside the region on every implementation
Outpainting Extends the image beyond its original boundaries Physically correct continuation of a scene
Regional prompting Applies different instructions to areas of a composition Sharp isolation when the implementation blends conditions
LoRA Learns a low-rank update for adaptation A universally best personalization method or permission to train on a subject
Keyframes and continuation Guide shot boundaries or extend a clip Exact frame continuity, consistent physics or available support on all models

Start with reference conditioning when it meets the requirement. Add adaptation only after an evaluation shows a persistent gap and the training rights are established. Keep adapter versions, base-model compatibility and permitted uses with the release. See LoRA, QLoRA and PEFT. ControlNet and reference adapters are described in their original ControlNet and IP-Adapter papers.

Production Pipeline Patterns

Start with the smallest workable design

A small image editor can send an authenticated request to one provider and return the resulting image if it fits the application's latency and timeout budget. Streaming previews may improve that interaction. Synchronous generation is not inherently invalid.

For longer work, separate acceptance from completion:

Architecture / visual model
flowchart LR U[Authenticated client] --> A[API: validate scope and reserve budget] A --> D[(Job database)] D --> Q[Durable work queue] Q --> W[Generation worker] W --> P[Provider or owned GPU service] P --> W W --> S[(Private output storage)] W --> D U --> R[Read status and authorized result] R --> D R --> S
Read diagram source
flowchart LR
    U[Authenticated client] --> A[API: validate scope and reserve budget]
    A --> D[(Job database)]
    D --> Q[Durable work queue]
    Q --> W[Generation worker]
    W --> P[Provider or owned GPU service]
    P --> W
    W --> S[(Private output storage)]
    W --> D
    U --> R[Read status and authorized result]
    R --> D
    R --> S

The queue may be fed through a transactional outbox: save the job and an enqueue record in one database transaction, then let a dispatcher publish the work. This avoids losing accepted jobs between a database commit and queue submission. A queue notification is permission to inspect a job, not permission to charge or execute it again.

Tip: state the distinction between an HTTP request timeout, a job deadline and a provider cancellation. They are different events.

Separate the identities

Identity Lifetime Purpose
Project and revision User's editable creative work Changing a script creates a new revision
Logical generation ID One requested generation or deliberate new variant An ordinary client retry retrieves the same job
Stage attempt ID One attempt to produce an intermediate output Investigate failed image, narration or encoding work
Provider job ID Provider's record of an accepted operation Reconcile completion, cancellation and billed usage
Asset ID and content hash One immutable stored output Reuse exact bytes and verify approval matches those bytes
Publication ID One approved delivery package Track the released files and their disclosures

Scope a client idempotency key to the authenticated tenant and operation. Save a canonical request hash with it. Reusing the key with different inputs returns a conflict; pressing “generate another variant” creates a new logical request. Do not treat identical prompts as a universal request to reuse the same image.

Model each stage as a recoverable operation

Architecture / visual model
stateDiagram-v2 [*] --> Queued Queued --> Submitting: lease and budget reservation Submitting --> Running: provider ID recorded Submitting --> OutcomeUnknown: response lost OutcomeUnknown --> Running: accepted job found OutcomeUnknown --> Queued: nonacceptance established OutcomeUnknown --> Held: cannot reconcile safely Running --> Validating: output copied and checked Running --> Failed: confirmed terminal failure Running --> CancelRequested: user cancels CancelRequested --> Cancelled: cancellation confirmed CancelRequested --> Quarantined: completion wins the race Validating --> Ready: required checks pass Validating --> Quarantined: invalid or prohibited output Ready --> [*] Failed --> [*] Cancelled --> [*] Quarantined --> [*] Held --> [*]
Read diagram source
stateDiagram-v2
    [*] --> Queued
    Queued --> Submitting: lease and budget reservation
    Submitting --> Running: provider ID recorded
    Submitting --> OutcomeUnknown: response lost
    OutcomeUnknown --> Running: accepted job found
    OutcomeUnknown --> Queued: nonacceptance established
    OutcomeUnknown --> Held: cannot reconcile safely
    Running --> Validating: output copied and checked
    Running --> Failed: confirmed terminal failure
    Running --> CancelRequested: user cancels
    CancelRequested --> Cancelled: cancellation confirmed
    CancelRequested --> Quarantined: completion wins the race
    Validating --> Ready: required checks pass
    Validating --> Quarantined: invalid or prohibited output
    Ready --> [*]
    Failed --> [*]
    Cancelled --> [*]
    Quarantined --> [*]
    Held --> [*]

This is a logical operation state machine, not a provider-specific API schema. Persist transitions with a version check so a stale worker cannot overwrite a newer state. Use renewable leases and stop dispatching when a worker loses its lease.

A provider may accept a costly job before the connection fails. Your database's idempotency key alone cannot prevent a duplicate provider charge. Use provider-supported deduplication within its documented scope and retention window, or query the accepted operation. If neither is possible, hold the ambiguous operation for reconciliation instead of blindly submitting again. See durable execution.

Verify webhook signatures where supported, reject replays that would change settled state, and deduplicate events. Treat callbacks as notifications; reconcile authoritative provider status where necessary. A periodic poller handles missed callbacks and jobs stuck past their deadline. An event arriving twice must not settle the same usage record twice.

Keep a private production manifest

Node editors such as ComfyUI express generation as connected operations and can save the workflow as JSON. Version that graph alongside model and custom-node dependencies. ComfyUI can also embed workflow data in generated files; inspect exported metadata before public delivery so private prompts and paths do not escape with an image. A visual graph still needs application-level permissions, budgets and durable recovery.

Store the information needed to explain a generation, subject to access and retention policy:

  1. Authenticated owner, project revision, purpose and approved usage scope.
  2. Input asset IDs and hashes, consent/rights references and prompt-template version.
  3. Encrypted prompt content where retention is permitted; never authentication secrets.
  4. Provider, requested and returned model versions, seed if supported, dimensions, duration, sampling settings and adapters.
  5. Workflow version, attempt IDs, provider IDs, timestamps and settled costs.
  6. Output hashes, evaluations, approval identity and final publication references.

A seed controls a source of randomness; it is not a complete reproducibility contract. Numerical kernels, hardware, precision, library versions and batching can affect results. Even local execution with fixed inputs needs explicit determinism controls, which can cost performance. Hosted reproducibility depends on the provider's documented contract. PyTorch reproducibility guidance describes these limits.

Preserving an approved asset is simpler than regenerating it exactly: retain its immutable bytes and manifest. Reproducibility is useful for debugging, but should not be the only way a customer can recover an approved file.

Control cost without breaking the product

Decision Benefit Cost or failure to manage
Reuse an authorized immutable asset Avoid regeneration and preserve the accepted result Recheck access, retention and permitted use; never share tenant-private cache entries
Generate cheap drafts first Reduce expensive renders for discarded ideas Draft and final quality can differ; validate the final rendition
Regenerate only an invalidated stage Preserve successful work Track dependencies; a changed narration can invalidate timing and lip-sync
Use a faster model or fewer steps Lower latency or cost where supported Measure prompt adherence and downstream rejection, not only call price
Batch compatible work Improve throughput or use provider discounts Longer waiting time; respect deadlines and tenant fairness
Keep workers warm Avoid repeated model-loading delays Pay idle capacity; reserve headroom for peaks
Bound retries and reserve spend Prevent runaway generation Some legitimate work waits or fails when its allowance is exhausted
Separate generation and encoding pools Scale the actual constrained stage More queues and operational complexity

Cache lookup uses tenant scope, immutable input hashes, model/workflow settings and the product's reuse policy. A cache hit is usable only while access and rights remain valid. Refresh an expired delivery URL for a stored asset; do not rerun the generator merely to get another URL.

For owned workers, size from arrival rate, measured service time, memory constraints and target utilization. Watch queue age, not just depth. For a hosted API, adding local workers does not increase provider concurrency or rate quotas. Apply backpressure before filling an unbounded queue.

Draft arithmetic: assume 100 requests, four candidate clips each, $0.80 per full render and $0.04 per draft. Rendering every candidate at full quality costs $320. Drafting all 400 and fully rendering 40 selected candidates costs $16 + $32 = $48, an 85% reduction under these assumptions. Selection rates, provider enhancement charges and rework can change that result. These are illustrative rates, not a quote for a named model.

Provenance and Safety

Separate origin, truth and permission

Question Evidence to collect Insufficient evidence
Where did this file come from? Validated creation/edit history and asset bindings A filename or a claim in the prompt
Has the bound content changed? Validate the signature and applicable content binding Merely displaying a credentials icon
Does it depict a true event? Independent factual evidence A valid provenance signature
May we use the subject or source material? Applicable rights, consent, license and usage scope Possession of an uploaded photograph
May we publish this rendition? Current policy checks and approval for the actual output Approval of an earlier draft

C2PA / Content Credentials is a standard for signed provenance records associated with digital assets. A valid record supports claims about the signing source and bound content; it does not establish that every assertion is true or that a photographed event happened. Its absence also does not prove that an asset is fake. See the C2PA explainer; the specification index currently lists version 2.4.

A hard binding uses cryptographic information to associate the manifest with specified asset content, following the format's binding rules. It need not mean a naive hash of every file byte: embedded manifests require appropriate exclusions. A soft binding, such as a watermark or fingerprint, can help discover associated provenance after some transformations. Recovery depends on the binding, transformation, detector and repository; it is not guaranteed to survive every edit. The C2PA technical specification defines these mechanisms.

For this design, keep the detailed production manifest private. Publish only appropriate provenance assertions and source references. Do not expose private prompts, customer IDs, credentials or source documents merely because a format supports metadata. Preserve provider credentials where possible and create a correctly linked new record for edited or transcoded derivatives. Protect signing keys and support revocation; a signed false statement remains false.

Watermarks and detection are additional evidence

A watermark embeds a detectable signal in content. Some systems are designed to tolerate common resizing, recompression or other edits. For example, SynthID covers several media types. It is not a universal detector for content produced by unrelated systems.

Robustness depends on the watermark and threat model. Research such as the NeurIPS 2024 watermark-removal study demonstrates attacks under specified conditions; it does not justify claiming every possible watermark always fails. Likewise, a classifier's “AI-generated” score is probabilistic evidence, not proof of origin or a substitute for consent records.

Put controls at input, output and publication boundaries

  1. Authenticate and authorize. Restrict source files, projects, collaborators and output access to their permitted scope.
  2. Validate uploaded files. Check decoded media, dimensions, duration, size and parser behavior. Use bounded processing and isolated converters.
  3. Establish rights and consent. Record permissions for voice cloning, likeness, music and brand assets with their intended use. A model license and a person's consent are separate requirements.
  4. Screen prompts and outputs. Apply the product's policies to sexual content, impersonation, violence and other prohibited uses. Evaluate false positives and missed violations; classifiers are fallible.
  5. Review consequential or ambiguous work. Route uncertain cases to qualified people and offer an appeal or correction path.
  6. Approve final bytes. Enforce exact revision, rendition hashes, policy version and unexpired authorization before release.
  7. Respond to abuse. Receive reports, preserve appropriate evidence, revoke delivery access and remove prohibited copies according to applicable obligations.

Laws depend on the system's role, jurisdiction, content and exceptions. Under EU Article 50, provider marking/detection duties and deployer deepfake disclosures are distinct obligations; standard editing and creative works have specific treatment. Article 50 must be read with the amended timetable: transparency applies from August 2, 2026, while qualifying systems already on the market have until December 2, 2026 for Article 50(2) marking/detection compliance. The Commission's current enforcement FAQ explains that transition.

In the US, the TAKE IT DOWN Act's platform provisions concern covered platforms and qualifying intimate imagery. The FTC describes a 48-hour removal duty after a valid request, with reasonable efforts to remove known identical copies. Do not generalize this into a blanket rule for every media service or every report. Use the FTC's compliance guide and the broader governance chapter when assigning requirements.

Evaluating Generative Quality

Quality is a collection of requirements. A beautiful image can contain the wrong product; an intelligible voice can say the wrong price. A preference vote cannot settle factual accuracy or ownership.

Use the right measure for the requirement

Measure What it measures Practical limit
FID: Fréchet Inception Distance Distance between Gaussian approximations to real and generated image feature distributions, conventionally using Inception features Depends on sample size, preprocessing, reference set and feature space; not a per-image truth score
CLIPScore Image-text compatibility using CLIP representations Can miss fine detail, counting, facts or context; alignment is not overall quality
FVD: Fréchet Video Distance Distributional distance in learned video-feature space Depends on temporal sampling, features and reference data; one low score does not certify a clip
FAD: Fréchet Audio Distance Distributional distance in audio-embedding space Embedding choice and reference distribution affect conclusions; not a speech-transcript correctness metric
MOS: mean opinion score Average ratings from listeners under a defined subjective test Specify the scale, question, listeners and conditions; scores from different protocols need not compare
OCR / transcription checks Text or speech content relative to required wording Recognizers can also err; verify critical facts directly
Human rubric and blind preference Task-specific quality judgments Requires representative cases, qualified raters and disagreement handling
VLM or audio-model judge Automated rubric assessment of media Can overlook defects or favor particular styles; calibrate against human review
pHash / SSIM or embedding similarity Particular forms of similarity between outputs A valid creative alternative may look different; similarity does not prove correctness

The original CLIPScore, FVD and FAD papers define different evaluation targets. Rethinking FID studies image-metric shortcomings and an alternative based on CLIP embeddings; FAD embedding research shows why encoder choice matters. No metric in this table establishes legal permission.

Build a release evaluation, not a beauty contest

  1. Version a test set covering actual content, languages, aspect ratios, difficult inputs and prohibited requests. Keep a held-out set for release decisions.
  2. Define hard constraints first: correct product identity, exact required text, supported format/duration, allowed source use and prohibited-content checks.
  3. Define quality rubrics: prompt adherence, artifacts, temporal consistency, speech intelligibility, lip synchronization and usefulness to the customer.
  4. Compare candidate and current systems on matched cases, with repeated generations when output variability matters. Blind and randomize subjective comparisons.
  5. Measure rejection, regeneration, review time and cost per accepted asset, alongside latency and preference.
  6. Apply practical regression margins and uncertainty intervals by important slice. One severe policy failure can block release regardless of an average score or statistical significance.
  7. Canary the change, watch drift and keep a rollback path for model, prompt, adapter and workflow versions.

A statistically significant improvement can be too small to matter. An important loss on a small language slice can lack statistical significance because the sample is insufficient. Report that uncertainty rather than calling it safe. Changes in input mix, judges or preprocessing can also cause score drift; do not immediately blame an unannounced provider update.

For byte-preserving operations such as retrieving a stored asset, exact equality is appropriate. For creative generation, evaluate requirements and distributions rather than requiring every new output to resemble a single “golden” picture. Public rankings are a candidate-selection aid; verify their current methods and test on your own distribution. See benchmarks and leaderboards.

The Model Landscape

Use this as a procurement checklist, not a ranking. Select from current API contracts and a task-specific evaluation, then pin the model or record the version returned by the service.

Requirement Current options to investigate Decision to verify
Hosted image generation/editing OpenAI's image guide lists gpt-image-2.5-sunburst; BFL provides FLUX image endpoints Reference support, edit fidelity, dimensions, version pinning, price and retention
Hosted multimodal video Google's video guide recommends Gemini Omni Flash for general generation; Veo 3.1 supports specified workflows such as extension and last-frame control The applicable API, supported inputs/durations, native audio and editing limitations
Another current video implementation BFL's FLUX 3 Video supports video generation/editing and optional native audio, with draft/enhancement workflows Supported duration/resolution, quota, actual draft-plus-enhancement billing and final quality
Owned image inference FLUX.2 klein 4B is one Apache-2.0 model; licenses differ for other variants The exact checkpoint license, hardware, safety requirements and adaptation compatibility
Speech and music Task-specific speech/TTS/music services or licensed source tracks Voice consent, music rights, allowed distribution, language quality and retention
Retired integration OpenAI Sora 2 models and the Videos API have a documented shutdown date of September 24, 2026 Migrate existing work; do not select the retired endpoint for a new service

Primary checks: OpenAI images, Google video generation, BFL FLUX 3 Video, FLUX.2 klein 4B model card, and OpenAI deprecations. The voice chapter covers speech stack selection.

Do not infer that an entire model family is open source or commercially usable from one permissive checkpoint. For example, BFL's non-commercial and self-hosted commercial terms differ, and the FLUX 3 page describes separate rollout stages for video, image, action and open weights. Model-use permission, output-use terms, source-material rights and likeness consent require separate checks.

An adapter must implement the actual provider contract. BFL's image quickstart returns a polling URL and a temporary output URL; it documents a ten-minute expiry for the signed result URL. Follow the documented polling destination after validating its provider origin, copy the output to authorized durable storage, and avoid giving customers a provider URL as their permanent asset record. Never send a provider credential to an arbitrary user-supplied URL.

Joint audio-video or a cascade?

Choice Good fit Tradeoff
Joint generation A short scene where motion, speech and ambient sound should be coordinated Changing one component may require regenerating more of the scene; evaluate synchronization rather than assuming it is perfect
Separate narration and video Exact approved wording, reusable voices, dubbing and independently edited tracks Must manage duration, transitions, loudness and synchronization explicitly
Hybrid Native ambient sound with a separately approved narration track More control, but careful mixing and prevention of conflicting speech are needed

In a cascade with visible speech, lip-sync depends on both the video and the final speech track. Putting lip-sync before its speech input is a dependency error. A voice-over on a product shot usually needs no lip-sync stage at all.

Interview design: a branded 30-second video service

Prompt: “Design a service where a business uploads a script and product images, reviews a draft, and downloads a narrated promotional video.”

1. Clarify scope

Ask about duration, required words, output formats, languages, identity/voice permissions, review responsibility and acceptable wait time. Confirm whether the service publishes to advertising accounts or only produces files.

For this interview, agree on the following scope: a private business workspace produces three-shot, 30-second videos with voice-over. It exports a final MP4 and captions. Publishing to social accounts, unrestricted celebrity cloning, live generation and long-form film editing are outside the first release.

2. Functional requirements

  1. Upload and authorize source pictures, script, brand settings and permitted narration voice.
  2. Validate inputs and display a cost estimate before paid generation begins.
  3. Create a storyboard and narration-text preview; let the customer approve or revise them before generating the media.
  4. Generate shots and compose the selected revision into a 30-second draft.
  5. Regenerate an individual shot without losing valid work from other stages.
  6. Show durable progress, failure reasons, cancellation state and actual usage.
  7. Check and approve the final rendition, then provide an authorized download and captions.
  8. Keep source, generation, approval and publication records; support retention and deletion policy.

3. Non-functional requirements

  1. Latency: acknowledge accepted work within one second at p95; target automated completion within ten minutes at p95 for the standard job class, excluding time awaiting customer review. Treat this as a target to validate under load.
  2. Durability: no acknowledged job is lost after a worker restart; status must explain partial completion or an unresolved outcome.
  3. Correctness: publish only the approved revision and validated rendition; preserve exact required words and product facts.
  4. Isolation: keep tenant inputs, outputs, cache entries and review records within authorized scope.
  5. Cost: enforce concurrent reservations, per-job limits and tenant limits before submitting additional billable work.
  6. Capacity: support 1,000 projects per working day, with four times average arrival rate at peak.
  7. Operability: expose stage latency, queue age, retry charges, rejection causes, unresolved provider operations and cost per accepted project.

Interview tip: a deadline is not a promise that every third-party request finishes. Say how the UI and refund/credit policy handle a deadline miss.

4. Estimate the scale

Assume a ten-hour active day, twenty working days per month, three ten-second generated shots per project, and a mean active generation time of 90 seconds per shot. Use a provider that supports the chosen clip contract; otherwise generate supported lengths and trim while accounting for their full charge.

Calculation Result Meaning
Projects per month 1,000 × 20 = 20,000 Denominator before rejection
Baseline shot jobs per day 1,000 × 3 = 3,000 Other stages need separate sizing
Peak shot arrival rate 3,000 / 36,000 × 4 ≈ 0.333/s Four times the active-day average
Peak rate with 15% additional generation attempts 0.333 × 1.15 ≈ 0.383/s Assumed rerender workload
Mean active slots at that rate 0.383 × 90 = 34.5 Arrival rate × mean service time
Slots at 70% target utilization ceil(34.5 / 0.70) = 50 A planning estimate, not a p95 queue guarantee
Encoded video at 8 Mb/s for 30 seconds 8 × 30 / 8 = 30 MB Before audio and container overhead
One month's final encoded videos 20,000 × 30 MB = 600 GB Excludes drafts, sources, replicas and downloads

For comparison, uncompressed 1920 × 1080 RGB frames at eight bits per channel, 30 frames/s, for 30 seconds occupy about 5.6 GB per video. Encoding changes the storage requirement dramatically. Do not estimate an encoded MP4 from its raw pixel count.

The 50 slots may be a provider concurrency allocation, not 50 physical GPUs. Owned hardware needs benchmarks for model, resolution, duration, batching, accelerator memory and co-location. Autoscaling cannot compensate for an unavailable provider quota.

5. Draw a baseline, then identify its flaws

The baseline is one worker that generates everything sequentially and saves the final file. It is enough to prove the product flow on a small workload.

Baseline flaw Observed consequence Repair and its cost
A worker owns all progress in memory Restart loses successful intermediate work Persist stage state and immutable assets; more storage and workflow logic
Retry the entire video on any failure Duplicate generation and different previously approved shots Retry only reconciled failed stages; requires explicit dependency tracking
One queue for model calls and encoding Long video calls block cheap tasks Separate stage queues and limits; additional scheduling
Approval refers only to project ID A revised or newly rendered file can bypass review Bind approval to revision and final asset hashes; repeated review when outputs change
Final link points to the provider Download expires or private inputs become accessible Copy to private storage and authorize delivery; storage/egress costs
One average “quality score” Wrong prices or rights violations pass a beauty test Hard constraints plus quality rubrics; more validation and qualified review

6. Develop the detailed design

Architecture / visual model
flowchart TB C[Workspace client] --> API[API: identity, scope, quotas] API --> DB[(Projects, jobs, approvals, cost ledger)] API --> UP[(Private source storage)] DB --> O[Outbox and durable orchestrator] O --> PLAN[Validate script and create storyboard] PLAN --> REVIEW[Approve storyboard and narration text] REVIEW --> IQ[Image queue] REVIEW --> AQ[Audio queue] IQ --> IMG[Image adapter] IMG --> VQ[Video queue] VQ --> VID[Video adapter] AQ --> TTS[Authorized TTS and licensed music] VID --> AS[(Immutable intermediate assets)] TTS --> AS AS --> COMPOSE[Composition and encoding pool] COMPOSE --> CHECK[Media, wording, policy and quality checks] CHECK --> FINAL[Final rendition approval] FINAL --> PUB[Publication transaction] PUB --> DELIVERY[Authorized CDN or signed download] O --> RECON[Webhook inbox and status reconciler] RECON --> DB CHECK --> DB
Read diagram source
flowchart TB
    C[Workspace client] --> API[API: identity, scope, quotas]
    API --> DB[(Projects, jobs, approvals, cost ledger)]
    API --> UP[(Private source storage)]
    DB --> O[Outbox and durable orchestrator]
    O --> PLAN[Validate script and create storyboard]
    PLAN --> REVIEW[Approve storyboard and narration text]
    REVIEW --> IQ[Image queue]
    REVIEW --> AQ[Audio queue]
    IQ --> IMG[Image adapter]
    IMG --> VQ[Video queue]
    VQ --> VID[Video adapter]
    AQ --> TTS[Authorized TTS and licensed music]
    VID --> AS[(Immutable intermediate assets)]
    TTS --> AS
    AS --> COMPOSE[Composition and encoding pool]
    COMPOSE --> CHECK[Media, wording, policy and quality checks]
    CHECK --> FINAL[Final rendition approval]
    FINAL --> PUB[Publication transaction]
    PUB --> DELIVERY[Authorized CDN or signed download]
    O --> RECON[Webhook inbox and status reconciler]
    RECON --> DB
    CHECK --> DB

The orchestrator creates versioned stage inputs. Before dispatch it checks the current project revision, scope, policy, budget and dependencies. Workers return immutable outputs; they do not decide which revision is public. Only the publication transaction changes the released package pointer.

The stage graph represents dependencies, so independent work can run concurrently:

Architecture / visual model
flowchart LR S[Approved script revision] --> N[Narration track] S --> B[Storyboard and product references] B --> V[Three generated shots] S --> M[Licensed or permitted music] N --> C[Compose voice-over video] V --> C M --> C C --> R[Encode required renditions] R --> Q[Check actual media and captions] Q --> A[Approve final hashes]
Read diagram source
flowchart LR
    S[Approved script revision] --> N[Narration track]
    S --> B[Storyboard and product references]
    B --> V[Three generated shots]
    S --> M[Licensed or permitted music]
    N --> C[Compose voice-over video]
    V --> C
    M --> C
    C --> R[Encode required renditions]
    R --> Q[Check actual media and captions]
    Q --> A[Approve final hashes]

If a later version includes a speaking avatar, insert a lip-sync stage after the relevant video and narration outputs. If a script change alters spoken duration, invalidate downstream timing and composition, and regenerate shots only when their content or duration requirements have changed.

Data model: projects, project_revisions, jobs, stage_attempts, assets, asset_dependencies, rights_records, approvals, usage_reservations, usage_settlements and publications. Unique constraints prevent duplicate logical jobs and settlements. Every lookup applies tenant scope. A versioned publication transaction checks required stage completion, asset hashes, current rights and approval before releasing the package.

API sketch:

Endpoint Contract
POST /projects/{id}/generations Authorized revision and idempotency key; return a stable job ID and reserved allowance
GET /jobs/{id} Current stage, safe status, cost summary and any action the customer must take
POST /jobs/{id}/cancel Request cancellation; report confirmed versus pending provider work
POST /projects/{id}/approvals Approve a specified revision/rendition set after access and role checks
POST /projects/{id}/publications Atomically validate and release the approved package
GET /assets/{id}/download Check current authorization and issue a short-lived download

A signed download URL is a temporary bearer capability. Expiry limits its duration; immediate revocation may need an authorization gateway or CDN invalidation strategy. Do not describe an already issued URL as instantly revoked merely because a database flag changed.

7. Walk through failures and repairs

Failure Response Remaining tradeoff
Third shot fails after two succeed Keep the first two; reconcile the failed attempt, then retry within budget A replacement shot may need continuity review
Provider accepted a request but response was lost Reconcile using provider identity/deduplication; hold an unknowable outcome Customer may wait; blind retry risks duplicate charges
Callback is forged, duplicated or late Verify, deduplicate, and apply versioned transitions Polling/reconciliation still needs capacity
Customer cancels while provider completes Stop new stages, quarantine late outputs, reconcile actual spend Cancellation may not reverse provider charges
Worker crashes after copying output Recover asset and stage records using stable IDs and hashes Clean up unreferenced uploads after a safe retention interval
Provider URL expires before copy Recover the same output through the provider if supported If unrecoverable, explain the failure before any paid regeneration
Final render changes the product label Fail the exact-content check and require correction/reapproval A cheap draft approval does not authorize a defective final
A customer loses permission to use a voice Block new use and apply the relevant removal/retention decision to existing assets Previously downloaded copies cannot be remotely erased
Provider outage or retirement Pause admissions or use a tested compatible fallback Different model output may require new approval and cost estimate
One tenant floods the service Tenant queues/limits, fair scheduling and global spend admission Some work is delayed rather than consuming all capacity

A fallback is a new implementation of the task, not a string substitution in an endpoint URL. Recheck allowed input types, licenses, output behavior, price, retention and evaluation thresholds.

8. Calculate complete operating cost

Assume these illustrative rates for 20,000 projects/month. They are deliberately separate from any vendor's current price card.

Cost Calculation Monthly estimate
Initial video generation 20,000 × 3 × 10 seconds × $0.08/s $48,000
Additional generation attempts 15% × initial generation cost $7,200
Storyboard images 20,000 × 4 × $0.03 $2,400
Narration and music allowance 20,000 × $0.05 $1,000
Composition and automated checks 20,000 × $0.10 $2,000
Storage and delivery allowance Assumed monthly amount $750
Maintenance and operations Assumed allocated monthly cost $4,000
Internal quality review 20% × 20,000 × 4 minutes / 60 × $45/hour $12,000
Total Sum of listed costs $77,350

This is $3.87 per attempted project. If 90% become accepted assets, the denominator is 18,000 and cost becomes $4.30 per accepted project. Internal review alone takes about 267 hours/month. Customer approval time is separate; music licensing, payment fees, taxes, support incidents or source retention may require additional line items.

If extra generation rises from 15% to 30%, add $7,200/month. A provider charging less per second can still be more expensive per accepted asset if its output triggers more rejection and review. Benchmark the full workflow before procurement.

9. Close the interview

“I would launch the scoped voice-over workflow with durable jobs, private immutable assets and approval of the actual final rendition. The main scaling controls are provider-aware admission and separate generation/encoding queues. The main correctness controls are revision-bound dependencies, reconciliation of unknown outcomes and a publication gate. I would validate the ten-minute target and unit economics with realistic traffic and rejection rates before expanding to avatars, more languages or direct social publishing.”

Interview Questions

1. Is every multimodal generation request necessarily asynchronous?

No. A short image operation can complete within the application's request budget. Long or interruption-prone jobs benefit from durable asynchronous execution. Explain the latency target, timeout behavior and recovery contract rather than prescribing one transport universally.

2. What is the difference between a diffusion model and a transformer?

Diffusion describes a generative modeling approach involving a noising process and learned reversal. A transformer is a neural-network architecture. A diffusion or flow-based generator can use a transformer; these are not exclusive categories.

3. Why does an idempotency key not automatically prevent two charges?

It deduplicates only where it is enforced. The local database can recognize the same request, but a provider may have accepted a timed-out submission. Safe recovery requires provider deduplication or reconciliation; otherwise the outcome remains uncertain.

4. The user asks for another image with the same prompt. Should the cache return the previous image?

Only if the product action requests reuse. “Generate another variant” is a new generation intent. Retrieval of an existing result is different from sampling again, even with identical text.

5. Can a fixed seed reproduce an approved video next month?

A seed alone cannot promise that. The execution environment, model version, parameters and provider contract matter. Store the approved bytes so access to the accepted result does not depend on regeneration.

6. How should a changed narration affect the workflow?

Create a new revision and invalidate dependent timing, lip-sync if present, composition and final approval. Reuse an unchanged shot only if its content, duration, rights and compatibility still satisfy the new revision.

7. Does a valid C2PA credential prove that an image is real?

No. It supports verification of signed provenance and bound content. Factual truth and permission to use the depicted subject require separate evidence. Missing credentials do not prove falsity either.

8. How do you scale a hosted video API?

Bound submissions to provider quotas, schedule tenants fairly, measure queue age and service time, and obtain additional capacity or test a fallback when needed. More local workers do not create more remote quota.

9. Why can the cheapest generator have the highest total cost?

Extra attempts, rejected work, internal review, encoding, storage and support affect the denominator. Compare complete cost per accepted asset at the required quality, not only the first-call rate.

10. Should you gate a creative-model release only on statistically significant FID improvement?

No. FID is one distributional measure. Use hard task requirements, safety constraints, representative human evaluation, practical regression margins, uncertainty and cost. Statistical significance does not establish usefulness or permission.

11. What if the draft was approved but the final render changes the text?

Final validation must inspect the delivered rendition. A changed label can fail an exact-content requirement even if the draft was correct. Repair it and obtain the required final approval.

12. How do you cancel a paid generation safely?

Record the request, stop dispatching new work and invoke provider cancellation if available. Reconcile whether generation completed and what was charged. A late result must not become public merely because the worker finished.

13. Why keep private production metadata separate from public provenance?

Debugging may require sensitive prompts, source references and account records. Public credentials should disclose appropriate origin/edit information without leaking private data. Both records can reference the same immutable asset identity.

14. When would you use a cascade instead of native audio-video generation?

When exact narration, separate language tracks or independent editing matter enough to justify synchronization work. Native generation can simplify coordinated scenes, but neither option universally has the best quality or lowest total cost.

15. What would you check before commercial use of open weights?

The exact model license, any commercial-use conditions, derivative/adapter terms, allowed deployment and output-use terms. Also check training/reference material rights and likeness or voice consent. A permissive software or model license does not grant those other rights.

Final revision cards

Remember Explain it in the interview
Intent → job → stage → asset → publication Different identities prevent retries, revisions and approvals from being confused
Accepted is not completed Durable status and reconciliation handle long-running provider work
Unknown is not failed Do not repeat a costly operation until retry safety is established
Approve the delivered bytes Draft approval alone cannot certify a changed final rendition
Origin ≠ truth ≠ permission Provenance, factual verification and usage rights answer different questions
Rate × time estimates active work Then add utilization headroom and test actual queue behavior
Cost / accepted assets Include rerenders, review, operations and rejected outputs
Model version is part of the workflow A provider change can require evaluation, migration and new approval

For practice, draw the workflow in five minutes, explain one ambiguous provider failure, then calculate how a doubled rejection rate changes cost and staffing. Close with the smallest useful release and the evidence needed to expand it.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Real-time voice agents: conversation, timing, and trustworthy actions
NEXT LESSONLearnastra AI Interview Guide →

Explore the diagram