ARTFEED — Contemporary Art Intelligence

Diffusion Transformers: Downscaled Training Gradients Analyzed

ai-technology · 2026-08-06

A recent paper on arXiv (2608.04448v1) explores the impact of training diffusion transformers with downscaled images on the preservation of the original training gradient signal. The authors analyze the gradient change due to downscaling, breaking it into two components: one influenced by noise and the downscale ratio, which diminishes at high noise levels as indicated by spectral analysis, and a second, constant floor linked to the absolute token count of the target grid, unaffected by noise. Their theoretical insights are supported by (route, sigma) maps, which reveal additional structures not evident in spectral analysis. For example, during the 1024->768 downscaling process, they pinpoint a range (0.65 < sigma < 0.95) that spectral criteria do not predict. The study also tackles the escalating training costs associated with higher resolutions, a significant hurdle in training diffusion models. The results indicate that while downscaling can maintain gradient signals at elevated noise levels, a fundamental limitation exists due to token count, affecting downscaling's efficacy at lower noise. This research enhances the understanding of the conditions under which downscaled training mirrors full-resolution training, with significant implications for the efficient training of high-resolution image generation models.

Key facts

  • Paper arXiv:2608.04448v1 analyzes gradient preservation in downscaled training of diffusion transformers.
  • Gradient change under downscaling decomposes into a noise-dependent term and a sigma-independent floor.
  • The noise-dependent term decays at high noise, consistent with spectral premises.
  • The sigma-independent floor is governed by the target grid's absolute token count.
  • Measured (route, sigma) maps corroborate the theoretical account.
  • On the 1024->768 route, a window (0.65 < sigma < 0.95) is identified that spectral analysis does not predict.
  • Training cost of diffusion transformers grows superlinearly with resolution.
  • The paper addresses unresolved questions about downscaled training gradient signal.

Entities

Institutions

  • arXiv

Sources