ARTFEED — Contemporary Art Intelligence

Motif 3: A 314B-Parameter Mixture-of-Experts Language Model

ai-technology · 2026-08-11

So, there's this new report on arXiv about a language model called Motif 3. It's a decoder-only model with a whopping 314 billion parameters, but only 13.2 billion are used for each token. The model features something called Grouped Differential Latent Attention, which combines different attention techniques. It also includes special connections and activations to enhance performance and efficiency. Motif 3 was trained on about 12.5 trillion tokens from various sources like websites, STEM fields, and different languages. It uses a unique structure with 384 experts in each layer, activating eight at a time. The report also touches on balancing experts and stabilizing numbers, though the abstract isn’t fully detailed. This model is a significant advancement in language technology!

Key facts

  • Motif 3 is a decoder-only Mixture-of-Experts language model.
  • It has 314 billion total parameters and 13.2 billion activated per token.
  • Each sparse MoE layer contains 384 routed experts, with eight selected per token.
  • The architecture uses Grouped Differential Latent Attention (GDLA).
  • It incorporates modified manifold-constrained hyper-connections, Expert Specific PolyNorm activations, and multi-token prediction.
  • Pretrained on approximately 12.5 trillion tokens.
  • Training data includes web documents, STEM, code, mathematics, multilingual content, and domain-specialized corpora.
  • The report is available on arXiv with ID 2608.09119.

Entities

Institutions

  • arXiv

Sources