Overview: Text diffusion models possess many advantages over autoregressive models for inference, test-time scaling, and RL. Recent work in converting strong autoregressive models into diffusion models shows the process is now cheap, such that any lab with a good AR checkpoint can adopt diffusion. As this process matures, I believe diffusion will rapidly gain popularity across the LLM ecosystem.

What’s diffusion?

Text diffusion differs from autoregressive generation in that you predict an entire sequence of tokens at once instead of just one at a time. The sequence begins as a canvas of random tokens. Predicting values for the entire canvas in one shot is difficult, so the model gets multiple attempts at it. At each step, it predicts every position and commits the positions it’s most confident about. The next step sees a slightly cleaner canvas, and so on until every position is committed.

Animation of block text diffusion: a grid of noisy tokens is denoised and committed over several steps, then appended to the prefix Animation of block text diffusion: a grid of noisy tokens is denoised and committed over several steps, then appended to the prefix

Figure 1: Block text diffusion. Each block starts as a canvas of noise. At every denoising step, the model commits the positions it is confident about (green). Once every position is committed, the block joins the clean prefix and a fresh canvas begins.

This allows us to generate tokens in only forward passes, where . Recent diffusion models use between 16 and 64 forward passes to generate 256 tokens. DiffusionGemma, for example, generates a 256-token canvas in at most 48 denoising steps, and averages about 12 with adaptive stopping.1 This isn’t free, since you have to make predictions for all tokens every forward pass.

Number passesMemory readsCompute FLOPs
Autoregressive
Diffusion

On balance this might not seem like a good tradeoff: you spend times more compute for at most savings in memory reads. But on GPUs, loading a byte from memory onto the compute die is hundreds of times slower than doing math with that byte. When decoding is bound by memory reads, the extra FLOPs are nearly free.

There are several variants of this basic form:

  • Repair: No tokens are committed until the very end, allowing each step to “overwrite” previously confident positions.
  • Self-conditioning: The model can use its previous guess to inform its next guess. This requires roughly twice the compute to train, but it is more data-efficient. In the appendix I extend previous works2 to train a novel variant that is more efficient and minimizes train/test mismatch.
  • Noise pattern: Instead of initializing the canvas with random tokens, you can initialize it with [MASK] special tokens. Recent work has argued that masking is suboptimal.345 Latent noise avoids materializing any tokens until the final step, keeping the canvas in latent or logit space throughout denoising.

The case from inference

Diffusion trades extra compute for fewer memory reads. Whether that trade pays off depends on where decoding sits on the roofline:

Memory-bound vs compute-bound decode regimes over batch size and sequence length Memory-bound vs compute-bound decode regimes over batch size and sequence length

Figure 2: Decode regimes over batch size and sequence length. Below the curve, decoding is compute bound; above it, decoding is memory bound, and diffusion’s extra FLOPs are nearly free.

Two things are loaded from memory at every decode step: the model weights, once per batch, and the KV cache, once per sequence. Large batches amortize the weight load and become compute bound. Long sequences push the other way, since every sequence brings its own KV cache. Past a certain length, the KV load alone takes longer than computing a token, and decoding is memory bound at any batch size.6

Low batch size and long context are valuable regimes in and of themselves. At 4,096 input tokens and 1,024 output tokens, DiffusionGemma is faster than its autoregressive counterpart (with MTP) up to a batch size of about 32.1 Note that this is an incredibly small context length. In thinking mode, DiffusionGemma’s own outputs average about 4,000 tokens, and GPQA-Diamond and LiveCodeBench traces run about 5,600–7,500. Longer sequences are more memory bound, which raises the batch size up to which diffusion wins.

MoE pushes in the same direction. Each token activates only a few experts, so compute per token shrinks while the weights loaded per batch stay large. Basically, the more memory bound we are, the more likely diffusion is a straightforward win over autoregressive decoding, even without speculative drafting.

Speculative decoding: One reason to suppose that diffusion will end up everywhere is because diffusion already is everywhere in the form of speculative decoding. DFlash and its successors7, which have been adopted by pretty much everyone8, use a diffusion draft model to predict an entire block of tokens in a single bidirectional pass. The diffusion draft model is a small distillation of the target model that has fairly similar output probabilities.

The main limitation is that verification is constrained to prefixes. The target is still a next-token predictor, so once it rejects a token, it must reject every token after it. DSpark9 and DFlash210 are largely a series of hacks to boost acceptance despite this constraint.

The main benefit of DFlash-style speculative decoding, which is to accept many tokens in a single forward pass in exchange for extra compute, is approximating what a diffusion model naturally does without needing to go through the bottleneck of an autoregressive verifier.

Beyond drafting: Diffusion models allow for number of techniques besides drafting to decrease the number of forward passes required for generation. The technique of pyramid sampling uses smaller models, often distilled from the target, in order to handle the first few denoising steps. Unlike in the autoregressive case, the confident positions don’t have to be biased toward the prefix, and can instead be spread throughout the canvas. Each model in the pyramid can even use its own decoding threshold.

Pyramid sampling is not lossless, since the output won’t match the target exactly. The same trade is well established in image diffusion, where a small model handles the noisy early part of the trajectory.11 It is also closely related to step distillation, where a student compresses several teacher denoising steps into a single transition.121314 Such a distillation is closely related to training a drafter for DFlash.

Lossless speculative decoding with a diffusion target is possible too, but drafting strategies built for autoregressive models need rethinking. One recent work uses samples of the next denoising step as speculations for several future steps, organizing these drafts as a calibrated graph.15 This is self-speculative (no separate draft model is needed) and cuts forward passes by up to 8.6× while preserving the output distribution. Related work extends speculative sampling to continuous diffusion16 and to whole denoising trajectories.17

Test-time scaling

Adaptive computation is natural with diffusion models. At inference time, an entropy threshold can be chosen based on the task. Lowering corresponds to committing fewer tokens per step, leading to more total steps for the whole canvas. By lowering we can spend arbitrary amounts of computation per canvas.18

We can compare diffusion to the increasingly popular architecture of looped or universal transformers. These loops seem to help with train time scaling1920, but are unable to provide out-of-distribution test-time scaling.21222324 Parcae, for example, sweeps the mean training loop count from 2 to 12 and finds that test-time gains saturate near that mean every time.20 Looping buys a compute knob over the range you trained on, not beyond it.

Why is this? First, consider KV caches. When you loop a transformer, every loop produces a new set of keys and values for every attention head. During training, each head learns to read the KV caches produced at the loop counts it sees. At test time, extra loops produce KV caches the head has never seen, and it can’t make sense of them. They’re totally out of distribution.

Secondly, you can consider the convergence of the residual stream. Iterations in a transformer need to converge in order to predict the next token. An attention head may be contractive, pulling the residual stream toward a fixed point, or it may be expansive, adding new context to the residual stream every loop. If the entire looped portion is contractive, the state converges to a fixed point where extra loops cannot be useful. If the loop is expansive, the state changes with every loop and the prediction diverges.

Chain of Thought, on the other hand, is a way to get arbitrary test-time scaling. For difficult problems, the model may learn (or can be prompted) to think for longer—that is, to emit more tokens—in a way that makes generating the correct answer more likely. CoT is an extremely elegant test-time mechanism. It appends to the prefix of the model such that the desired output is more likely when conditioned on this prefix than on the original prompt. Progress lives in a separate state and is attended to in the KV cache. Each attention head is already very good at digesting a prefix of any length, so CoT can be trained without mismatch between training and inference.

The situation is even more elegant with text diffusion. The KV cache is identical during training and generation: a clean prefix and a noisy canvas. During training, the noise level of the canvas is sampled across its whole range, so every state the model reaches at inference is one it could have seen in training. Finer sampling gives denser coverage, but even coarse sampling spans a massive range of states. Looping has no such guarantee, since extra loops produce states that training never covers. No attention head’s job changes from train time to test time. And unlike looping, what contracts across denoising steps is the canvas itself, so we can keep adding denoising steps and expect the canvas to cohere.

In practice, diffusion shows test-time scaling out of the box. As DiffusionGemma takes more denoising steps, its score climbs from about 9 to 59:

DiffusionGemma average score vs effective denoising steps after SFT DiffusionGemma average score vs effective denoising steps after SFT

Figure 3: Average score on GPQA-Diamond and LiveCodeBench-v6 of DiffusionGemma after SFT (but prior to SD·RL), versus effective denoising steps. Data reproduced from the technical report.1

This is true test-time scaling. During training, the model never sees more than one or two denoising iterations at a time, so every additional step at inference is, to some extent, “new.” Yet the model makes use of it immediately.

Self-Correction: One of the major reasons Chain of Thought works is self-correction. The model can propose an answer, notice that it’s wrong, and correct it.2526 Repair plays a similar role in diffusion. When committed tokens can be overwritten, a later denoising step can reject an earlier proposal.27 Self-conditioning helps in the same way, since each denoising step sees the model’s previous guess for every position. In the appendix I describe a self-conditioning recipe efficient enough to make this extremely useful.

Diffusion also scales with parallel compute. Sequential Monte Carlo (SMC) runs many drafts side by side and resamples toward the promising ones at every denoising step, without waiting for the final answer.28 Self-rewarding SMC does this with no reward model at all, using the model’s own confidence along each trajectory as the weight.29

Diffusion’s test-time compute has a natural advantage over CoT, in that the KV cache doesn’t grow with additional computational steps. With diffusion, once the thinking is paid for and the tokens are committed, you do not need to continue paying for (or attending to) that effort in future blocks. This is a bit closer to how humans think. While writing this post I went from a bare skeleton to a rough draft to a final essay. Getting rid of previous versions allows me to have a clear working space. If my iteration process was closer to CoT, I’d have to continually attend to the entire history of work between subsequent drafts. That said, the two methods are totally orthogonal, and diffusion pairs with reasoning as a separate axis for scaling test-time compute.

RL for diffusion

Thinking with trajectories: Early attempts at RL on diffusion models had a lot of problems to work out. Depending on how you parameterize the diffusion model, it’s not immediately clear that you can even estimate the logprobs of a trajectory, or whether conditioning on the noise biases the estimate. Some prior literature estimates sequence likelihoods with the ELBO,30 which gives noisy importance ratios, and then works to reduce that noise or keep training from collapsing.31323334 Other work uses a one-step mean-field proxy, which is cheap but biased.3536 As I’ll summarize here, these problems have been largely worked out. The difference from autoregressive is that you have to factor the sampling probabilities in terms of trajectory states.

Consider estimating the probability of a generation conditioned on a set of initial noise . (If you renoise between every step, you can imagine is a frozen tape that you read off after each round).

Here, and represent the canvas and the committed set of tokens, respectively, at step . You can estimate in an unbiased way as an expectation over both and :

To compute this, we need the model’s prediction at every committed position, at the step where it was committed. So per block we generate, we have to save all the intermediate canvases from the rollout. Each is a separate forward pass over the canvas, so a block with denoising steps costs forward and backward passes.

Subsampling: The innovation that makes this efficient is subsampling. Instead of computing a backprop for each denoising state, we can pick a random subset of steps and rescale by , and the estimate stays unbiased:

The analogue for autoregressive policy optimization would be only sampling some of the token gradients. While this is also unbiased, it’s not an efficient tradeoff, since we still have to do basically the same number of calculations.

There are a couple of cases where this breaks. In diffusion-with-repair, committed tokens can be overwritten, the committed set is no longer monotonic, and the factorization breaks, since many different trajectories can end at the same . The easiest solution is just to ignore the factorization and score the whole trajectory instead. The reward only depends on , which is a function of the trajectory, so REINFORCE over the trajectory is still unbiased, just noisier.

The other case is self-conditioning. Each prediction now depends on the model’s own guess from step , which depends on the guess from step , and so on back to the start of the block. To account for this correctly you’d need to backprop through the whole chain, which is not efficiently batchable. One fix is to freeze the recurrent depth at 1 or 2, saving each guess alongside its canvas during the rollout, and backprop only through one or two steps. This is still fairly efficient, especially with subsampling, but the gradient is now slightly biased.

Outside of these cases, the importance ratios can be calculated exactly, computed per token at the step it was committed, just like in the autoregressive case.

On-policy: Diffusion should also make rollouts faster. If we want more on-policy steps, we have to pay for this with lower batch sizes, since every rollout has to come from the current policy. This is exactly the regime where we should expect diffusion to be faster than AR. Rollouts also tend to be long, where reasoning traces run to thousands of tokens, which pushes decoding even further into the memory-bound regime.

Dense rewards: Denoising is likely a better way to decompose rewards along a trajectory than AR. In AR we’re constrained to consecutive token prefixes for reward assignment, but diffusion allows us to score entire intermediate states. The change in reward between consecutive drafts is concretely interpretable as when the answer was generated, and therefore how to more correctly allocate credit.37 Early drafts may be reward-sparse, especially when there is an ultimate verifiable reward, like with checked final answers or unit tests. On-policy self-distillation (OPSD) is one way to densify this signal, and it’s much more natural with diffusion—in AR it’s hard to avoid having the teacher leak causality, while a diffusion teacher already sees the entire draft at once.

Why now?

In this post I’ve summarized some advantages of diffusion over autoregressive models, but many of these points are not fundamentally new. Why should we expect diffusion to suddenly become popular? As mentioned earlier in the post, distilled diffusion models already are everywhere, in the form of draft models. This technique saw widespread adoption this year because well-written papers contained simple recipes. I believe the release and report of DiffusionGemma will play a similar role but for the technique of converting entire fully-trained models into diffusion variants.

DiffusionGemma is a blueprint: Converting an autoregressive model into a diffusion model goes back nearly to the start of text diffusion.3839404142 What’s impressive about DiffusionGemma is how cheap the conversion has become. It skips pretraining entirely and warm-starts from the final post-trained Gemma 4 checkpoint. After a short SFT stage on comparatively little data, the RL stage runs for only about 1,400 steps.43 The whole recipe uses under 10% of the AR model’s training tokens.1 The result pays a moderate quality cost relative to its AR initialization in exchange for a 5–7× speedup. It also vets a series of modern diffusion tricks, including block-wise canvases, self-conditioning, and entropy-bounded sampling.

Live research threads: Diffusion also sits at the intersection of many of the most interesting research threads right now. In this post I’ve tried to motivate diffusion’s advantages for inference optimization, test-time compute, and dense-reward RL. Structured generation, as seen in Jev44 and OpenAI’s recent Decisions API,45 is another major thread. A recent vLLM PR makes minor changes to instantly turn DiffusionGemma into a structured decision engine.46 Data-sparse regimes, another popular thread, seem to benefit from diffusion-based pretraining, since seeing every example under many noise patterns acts as a strong form of data augmentation.47

The future frontier: What does the future of discrete diffusion look like? Much faster self-hosted models, real-time compute, arbitrary-length canvases, and easy structured generation seem to be in the near future. I predict that diffusion models will become anti-modal, not constrained to any single modality, whether text, audio, image, or video. As inference becomes cheaper, byte-level diffusion may even become attractive. Many of the greedy choices baked into today’s tokenizers were made for autoregressive models, and they may carry different tradeoffs for a world of diffusion.

Appendix: self-conditioning

Self-conditioning lets the model see its own previous guess.2 Text diffusion models commonly feed back the predicted token distribution, often as the probability-weighted average of the token embeddings, .4849

Instead of the distribution, I propose feeding back the final hidden state. Let be the last layer’s output at a noised position. On the next pass, add a learned projection of to the token embedding at the same position:

The latent, the logits, and the probabilities carry the same information about the prediction (barring some details about temperature), but the latent is the simplest and cheapest form. The temperature can be added as an extra input if needed.

The recipe in Analog Bits2 runs both passes on the same noisy canvas. At sampling time, though, the self-conditioning input always comes from the previous denoising step, which saw a noisier canvas. In between, the sampler committed some tokens. So training and sampling use self-conditioning in different situations.

Reveal training matches the sampler. For each block:

  1. Draw , then , so .
  2. Draw one uniform per position. Pass 1 (no gradient) sees , where position is noised if .
  3. Pass 2 (trained) sees , where position is noised if , with fresh noise tokens. Since , the noised positions of are a subset of those of ; the rest are revealed, set back to the truth. Marginally, each position of is noised with probability , the same as in single-pass training.
  4. A fraction of blocks receive pass 1’s latent as self-conditioning. The original Analog Bits paper proposes ; I find works better, though results are fairly insensitive to this hyperparameter.

Each position goes through one of four transitions between the two passes. Pass 1 stands in for the previous denoising step, and pass 2 for the current one:

  • Already committed (): the token stays, and the feedback was computed with the token present.
  • Newly committed (): a token now sits in the slot, and the feedback is from when it was noise.
  • Still uncommitted (): the slot gets fresh noise, and the feedback is from a different noise token.
  • Un-committed (): only happens with repair.

The Analog Bits recipe only trains the diagonal, producing cases at generation time that were never seen during training. Revealing and renoising between the two passes covers the off-diagonal case, which is exactly the prediction–canvas mismatch that shows up during generation.


def forward(prefix, canvas, h_prev=None):
    x = embed(canvas)
    if h_prev is not None:
        x = x + W_sc(rmsnorm(h_prev))         # latent self-conditioning, W_sc is zero-initialized
    h = rmsnorm(transformer(rmsnorm(x), context=prefix))   # canvas attends to the clean prefix and itself
    return lm_head(h), h
 
def train_step(prefix, x0, q=0.75):           # x0: the clean block, P tokens
    t2 = uniform(0, 1)
    t1 = uniform(t2, 1)                       # t1 >= t2
    u = uniform(0, 1, size=P)                 # one draw per position
    z1 = where(u < t1, random_tokens(P), x0)  # pass 1 canvas (noisier)
    z2 = where(u < t2, random_tokens(P), x0)  # pass 2 canvas, noise is a subset of z1's, fresh tokens
    with no_grad():
        _, h_prev = forward(prefix, z1)
    if uniform(0, 1) > q:
        h_prev = zeros_like(h_prev)           # some blocks train without self-conditioning
    logits, _ = forward(prefix, z2, h_prev)
    return cross_entropy(logits, x0)
 
def entropy_bound(H, available, tau):
    # lowest entropy first; keep taking positions while the sum of the earlier ones stays <= tau,
    # so every step commits at least one token
    order = argsort(where(available, H, inf))
    H_sorted = H[order]
    take = (cumsum(H_sorted) - H_sorted <= tau) & available[order]
    return scatter(order, take)
 
def sample_block(prefix, tau):
    canvas = random_tokens(P)
    committed = zeros(P, dtype=bool)
    h_prev = None
    while not committed.all():
        logits, h_prev = forward(prefix, canvas, h_prev)
        accept = entropy_bound(entropy(logits), available=~committed, tau=tau)
        canvas = where(accept, sample(logits), canvas)
        committed |= accept
        canvas = where(committed, canvas, random_tokens(P))   # renoise everything uncommitted
    return canvas

Algorithm 1: Latent self-conditioning with reveal training, and the matching entropy-bounded sampler (one block, no repair).


If this post was useful to you, you may cite

@misc{
      srivastava2026,
      author = {Varun Srivastava},
      title = {Diffusion will be everywhere},
      year = {2026},
      url = {https://varunneal.github.io/essays/diffusion}
}

Footnotes

  1. Google DeepMind 2026 DiffusionGemma Technical Report ^ ^ ^ ^

  2. Chen, Zhang, Hinton 2022 Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning ^ ^ ^

  3. Chen, Wang 2026 Parallelism, critical windows, and separations among diffusion language models ^

  4. Sahoo et al 2025 The Diffusion Duality ^

  5. Schiff et al 2024 Simple Guidance Mechanisms for Discrete Diffusion Models ^

  6. Back of the envelope equation for this graph: consider a decode step with batch size , KV length , and parameters, on hardware that does FLOPs per byte loaded. Up to constants, memory time scales as : the weights load once per batch and the KV cache once per sequence. Compute time scales as . I ignore attention, since its compute is always hidden under its own KV load. Decoding is compute bound when , i.e. below the curve . The curve starts at and saturates at , where loading one sequence’s KV cache takes as long as computing one token. That asymptote may only be reached at absurdly long sequences, but the general shape is right. Linear attention changes the picture entirely, though as long as there is at least one full-context attention layer, the diagram stays broadly accurate. ^

  7. Chen, Liang, Liu 2026 DFlash: Block Diffusion for Flash Speculative Decoding ^

  8. DFlash runs in SGLang, vLLM, TensorRT-LLM, and llama.cpp. Meta, Poolside, Xiaomi, and NVIDIA ship official DFlash drafters with their own models, and Modal “use[s] it with every compatible model.” The evidence is so overwhelming that no one even talks about it anymore. See the DFlash 2 announcement for a fuller list. ^

  9. DeepSeek, Cheng et al 2026 DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation. Adds a lightweight autoregressive pass that takes into account the transition between tokens, alongside many other effective tricks. ^

  10. Inco AI 2026 DFlash 2 ^

  11. Pan et al 2024 T-Stitch: Accelerating Sampling in Pre-Trained Diffusion Models with Trajectory Stitching ^

  12. Salimans, Ho 2022 Progressive Distillation for Fast Sampling of Diffusion Models ^

  13. Lu et al 2026 On Policy Transition Distillation for Diffusion Language Models ^

  14. Dat, Li, Wang 2026 dOPSD: On-Policy Self-Distillation for Diffusion Language Models ^

  15. Agrawal et al 2025 Structuring the Future: Diffusion LLM Speculative Decoding via Calibrated Draft Graphs ^

  16. De Bortoli et al 2025 Accelerated Diffusion Models via Speculative Sampling ^

  17. Pan et al 2026 Trajectory-Level Speculative Decoding for Diffusion Language Models. Allows the target to deviate from the draft trajectory, at the cost of introducing bias. ^

  18. E.g. via repair, or with samplers that don’t require committing any tokens per denoising step. ^

  19. Chen, Vegesna, Dahal, Wilson 2026 How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents ^

  20. Prairie et al 2026 Parcae: Scaling Laws For Stable Looped Language Models ^ ^

  21. Dehghani et al 2018 Universal Transformers ^

  22. Geiping et al 2025 Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach ^

  23. Saunshi et al 2025 Reasoning with Latent Thoughts: On the Power of Looped Transformers ^

  24. Jolicoeur-Martineau 2025 Less is More: Recursive Reasoning with Tiny Networks ^

  25. Lilian Weng 2025 Why We Think ^

  26. Kumar et al 2024 Training Language Models to Self-Correct via Reinforcement Learning ^

  27. Wang et al 2025 Remasking Discrete Diffusion Models with Inference-Time Scaling ^

  28. Yadala Chanchu, Abdulsamad, Naesseth 2026 Discrete Diffusion Inference-Time Control with Nested Sequential Monte Carlo ^

  29. Luo et al 2026 Self-Rewarding Sequential Monte Carlo for Masked Diffusion Language Models ^

  30. The ELBO (evidence lower bound) is a lower bound on the marginal likelihood , the probability of summed over every initial noise pattern and denoising trajectory. That marginal is intractable, which is why these methods have to estimate it. As I’ll argue below, RL never needs it, since the likelihood of the trajectory the sampler actually took is enough. ^

  31. Zhu et al 2025 LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models ^

  32. Rojas et al 2025 Improving Reasoning for Diffusion Language Models via Group Diffusion Policy Optimization ^

  33. Ou et al 2025 Principled RL for Diffusion LLMs Emerges from a Sequence-Level Perspective ^

  34. Zhong et al 2026 Stabilizing Reinforcement Learning for Diffusion Language Models ^

  35. Zhao et al 2025 d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning ^

  36. Tang et al 2025 wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models ^

  37. Writing for the model’s full guess at step , rewarding the difference is potential-based reward shaping. The per-step rewards telescope to the final reward, so the optimal policy is unchanged. Ng, Harada, Russell 1999 Policy Invariance Under Reward Transformations ^

  38. Gong et al 2024 Scaling Diffusion Language Models via Adaptation from Autoregressive Models ^

  39. Ye et al 2025 Dream 7B: Diffusion Large Language Models ^

  40. Cheng et al 2025 SDAR: A Synergistic Diffusion-AutoRegression Paradigm for Scalable Sequence Generation ^

  41. Radical Numerics 2025 RND1: Simple, Scalable AR-to-Diffusion Conversion ^

  42. Bie et al 2025 LLaDA2.0: Scaling Up Diffusion Language Models to 100B ^

  43. Estimated, based on the x-axis of the SD·RL training curves, which are plotted in buckets of 200 steps. ^

  44. TypeSafe AI’s Jev returns typed decisions, such as a choice from a fixed answer set, with calibrated confidence. See this writeup. ^

  45. OpenAI 2026 DevDay 2026 recap. See also this practical guide. ^

  46. Mastracci 2026 “We have Jev at home”. Uses canvas infilling to automatically turn DiffusionGemma into a structured decision engine. ^

  47. Prabhudesai et al 2025 Diffusion Beats Autoregressive in Data-Constrained Settings ^

  48. Dieleman et al 2022 Continuous diffusion for categorical data ^

  49. Strudel et al 2022 Self-conditioned Embedding Diffusion for Text Generation ^