ARTFEED — Contemporary Art Intelligence

Meta Doubles GEM Training Efficiency to 20-25% MFU with Custom Kernels and 5D Parallelism

ai-technology · 2026-08-03

Meta has unveiled its Generative Ads Recommendation Model (GEM), which now operates at LLM scale using thousands of GPUs. Over the past year, it has achieved a Model FLOPs Utilization (MFU) of 20-25% and quadrupled its training FLOPs. This success stems from collaborative efforts in optimizing kernels, precision, parallelism, networking, and memory. GEM’s design integrates trillions of sparse parameters alongside billions of dense ones, utilizing ad content and user engagement data for training. A specialized kernel library was created, featuring Jagged Flash Attention (JFA) and Generalized Dot-Product Attention (GDPA), along with topology-aware 5D parallelism. These advancements resulted in a 40-140% increase in TFLOPS from JFA v4, a twofold speed enhancement from GDPA, and a 30.6% MFU boost from BlockAttention.

Key facts

  • Meta's GEM model now trains at LLM scale on several thousand latest-generation GPUs.
  • E2E training efficiency doubled to 20-25% MFU.
  • Training FLOPs scaled 4x in 12 months.
  • Custom kernel library includes JFA, GDPA, and BlockAttention.
  • Mixed ultra-low-precision training with MXFP8 attention and MLP.
  • Topology-aware 5D parallelism with SM-free collectives.
  • JFA v4 achieves 40-140% TFLOPS improvement over JFA v2.
  • GDPA kernel achieves 2x forward speedup and up to 3.5x over Flash Attention 4.
  • BlockAttention improves self-attention layer MFU by +30.6% over Triton block attention.
  • Base Batch Shuffling (BBS) reduces load imbalance with zero cross-rank communication.

Entities

Institutions

  • Meta
  • Instagram
  • Facebook
  • PyTorch
  • NVIDIA

Sources