A diffusion language model generates text through a learned denoising process. In a common masked-discrete formulation, training corrupts token sequences by masking positions and teaches a model to recover missing tokens. Generation starts from masked positions and iteratively resolves them. Some systems operate over whole sequences; others generate blocks while retaining an autoregressive order between blocks.
This chapter explains the mechanism and the deployment decision. “Diffusion” alone does not establish faster responses, correct reasoning, or a particular API capability.
How They Work
Contrast two factorization and generation choices:
| Autoregressive generation | Masked diffusion generation |
|---|---|
| Predict the next token from a preceding prefix | Predict masked positions using available context |
| Newly generated tokens create a sequential dependency | Several positions may be resolved during one iteration |
| Conventional KV reuse relies on unchanged causal history | Changing bidirectional state complicates straightforward cache reuse |
| Streaming naturally follows the committed prefix | Partial output may require revisions or block commitment |
In a simplified masked process:
- During training, choose a corruption level and mask selected clean tokens.
- Train the predictor to recover masked content with the method's specified objective.
- During generation, start with masked output positions and fixed conditioning input.
- Predict candidate content and select positions to commit or revisit according to the sampling schedule.
- Repeat until the sequence or block reaches its completion rule.
LLaDA is a primary example of masked language diffusion. Denoising schedules differ; do not say every model uses the same confidence rule, remasking behavior, or fixed number of steps. A diffusion likelihood-bound objective is different from ordinary autoregressive likelihood, but that fact alone does not prove a universal quality ceiling.
Read diagram source
flowchart LR
A[Prompt and masked output block] --> B[Predict missing positions]
B --> C[Apply schedule and commitment rules]
C --> D{Block complete?}
D -->|No| B
D -->|Yes| E[Commit output or advance to next block]
The Speed Advantage and the Tradeoff
The opportunity is to commit several useful tokens per costly model computation. Additional denoising iterations, attention over a block, preparation, and output handling still cost time. More aggressive commitment may hurt quality; more steps may recover quality but reduce the latency benefit. Neither relationship is an exact universal curve.
Worked comparison: an illustrative AR service emits 256 tokens in 256 steps averaging 8 ms, for about 2.048 seconds of generation. A block model taking 16 iterations at 40 ms uses 0.640 seconds; at 64 iterations it uses 2.560 seconds. Add prefill, queueing, and network time to both. The faster configuration is useful only if it meets the same quality and output requirements.
Distinguish generation tokens/second from first useful output, complete-response latency, aggregate throughput, and energy. A model that revises a block may need a different user-interface and streaming contract from a simple append-only token stream.
The 2026 Landscape
The following is a September 2026 capability snapshot, not a speed ranking:
| Family or offering | What public primary material establishes | What to verify for a product |
|---|---|---|
| Mercury | Commercial diffusion-model API offering | Current model identifier, schema/tool support, rate limits, pricing and service terms |
| Gemini Diffusion | Experimental text-diffusion demo | Do not assume a demo is a supported production API |
| DiffusionGemma | Open-weight block generation using an encoder/cached context and a denoising decoder | Exact artifact, modalities, runtime, context and memory behavior |
| LLaDA | Open research on masked diffusion language modeling | Check the specific successor checkpoint and serving implementation |
| Dream 7B | Research and released models with diffusion, infilling and schedule flexibility | Task quality and schedule cost under the deployed configuration |
Keep vendor speed and quality claims attached to their benchmark conditions. A broad claim that no diffusion model can do a certain task, or that all diffusion models dominate code generation, is too strong.
Hybrids: Draft with Diffusion, Verify with AR
There are several ways to combine parallel refinement and sequential commitment:
- Block diffusion: autoregressive order across blocks and denoising within each block. Block Diffusion studies this combination. Block size affects parallelism, caching, and output behavior.
- Diffusion proposals with AR selection: propose several positions in parallel and use an autoregressive commitment mechanism. TiDAR describes a single-model hybrid using structured attention. Exact-distribution claims require the particular sampling proof, not merely the word “hybrid.”
- Adapting an AR checkpoint: Fast-dLLM v2 investigates block-diffusion adaptation and hierarchical caching. Reuse of a checkpoint still requires training and evaluation.
- Post-training for reasoning: d1 studies masked SFT and a diffusion-specific policy-gradient method. Ordinary autoregressive RL code is not automatically compatible with a different generation process.
These are mechanisms to understand, not a forecast that one architecture must replace every other one. Relate them to speculative decoding and RLVR without conflating their guarantees.
Choose a pilot and an evaluation
For a code-editing assistant, infilling and block refinement may be useful because the desired change is not always a left-to-right continuation. Start with a contained set of edits and a working AR baseline.
| Requirement | Evaluation | Failure that changes the decision |
|---|---|---|
| Correct edit | Tests, diff review, preserved surrounding behavior | Fast output that breaks unrelated code |
| Responsive interaction | Time to first usable edit and full completion | Large commitment delay despite high reported tokens/second |
| Structured tool calls | Schema and semantic validity | Unsupported or malformed actions |
| Long context | Evidence-use tests at target lengths | Relevant constraints ignored or cache cost excessive |
| Stable streaming | Actual client protocol behavior | Revisions misinterpreted as committed text |
| Sustainable operation | Cost, load, version support and rollback | Benchmark win disappears under concurrency |
Maturity and What to Do Today
Define the user task, baseline quality, latency target, and allowed operational complexity. Then compare implementations that satisfy the same contract. A specialized pilot may justify an additional model, but its routing, monitoring, and maintenance costs belong in the decision. Preserve a tested fallback and avoid changing architecture solely because of a headline speed result.
Interview practice
- What is being denoised in a masked text model? Corrupted token positions, often represented by mask tokens, under the specified training and sampling process.
- Does parallel prediction mean one-step generation? No. Several denoising iterations may be needed before content is committed.
- Why is cache reuse harder with changing bidirectional context? Earlier representations can change when other positions change; ordinary causal cache assumptions may no longer apply.
- How do blocks help? They permit a defined history between blocks while allowing parallel work within a block, with architecture-specific caching.
- Can a diffusion model make tool calls? Check the specific model and interface. The architecture name alone proves neither support nor impossibility.
- What benchmark would convince you to adopt it? Matched task quality and valid outputs, better user-visible latency or cost under load, and acceptable operational support.
Recall card and closing
Corrupt → predict → refine → commit → evaluate. Explain the commitment and cache behavior before citing a speed. Close with the workload that benefits and the conditions under which the baseline remains preferable.