ARTFEED — Contemporary Art Intelligence

xPress: A Causal Refiner for Diffusion Drafters in Speculative Decoding

ai-technology · 2026-08-04

A recent study, identified as arXiv paper 2608.02438, presents xPress, a novel approach aimed at enhancing the efficiency of speculative decoding with diffusion drafters. This research tackles a significant drawback found in block-diffusion drafters like dFlash, which create an entire block of draft tokens in one forward pass. Although this method minimizes the overhead associated with drafting multiple tokens, the final phase of the single-pass discrete denoising process samples tokens that are conditionally independent based on logit distributions for each position. As a result, the drafts consist of marginals rather than a joint distribution, causing generated sequences to include tokens that, while individually probable, are collectively unlikely according to the target model's distribution, which can lead to early rejections and restrict acceptance length. To address this issue, xPress is introduced as a lightweight causal refiner that reinstates the necessary causality in diffusion drafters. The full paper can be accessed at https://arxiv.org/abs/2608.02438.

Key facts

  • Paper arXiv:2608.02438 introduces xPress.
  • xPress is a lightweight causal refiner for diffusion drafters.
  • Block-diffusion drafters like dFlash generate entire blocks of draft tokens in a single forward pass.
  • The method addresses the issue of independently sampled marginals in speculative decoding.
  • xPress aims to restore causality in diffusion drafters.
  • The paper is available on arXiv.
  • The approach reduces overhead of multiple-token drafting.
  • The method targets early rejection and limited acceptance length.

Entities

Institutions

  • arXiv

Sources