ARTFEED — Contemporary Art Intelligence

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

ai-technology · 2026-08-06

A recent paper on arXiv (2608.03457) introduces LLaDA MoE v2, which systematically examines the scaling behavior of Mixture-of-Experts (MoE) diffusion language models (dLLMs). This study reveals notable quantitative distinctions from autoregressive (AR) models concerning optimization, compute distribution, and architecture. Among the significant results: the optimal nominal batch size increases more rapidly with compute, the optimal learning rate diminishes quickly, and IsoFLOP analysis indicates a data-side bias, with the optimal token budget expanding faster than the computation activated on the model side. For MoE architecture, larger scales benefit from bigger expert pools at a constant activated capacity, while moderate expert granularity is still effective, and the optimal fraction of activated capacity for shared experts remains consistent. The paper establishes new scaling laws for MoE dLLMs, providing valuable insights for future model development.

Key facts

  • Paper title: LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
  • arXiv ID: 2608.03457
  • Announce type: new
  • Focus: scaling behavior of MoE diffusion language models
  • Optimal nominal batch size grows faster with compute
  • Optimal learning rate decays more rapidly with compute
  • IsoFLOP analysis reveals data-side tilt in compute allocation
  • Larger scales favor larger expert pools at fixed activated capacity

Entities

Institutions

  • arXiv

Sources