MOSAIC: Systems-Aware Scaling for Sparse Mixture-of-Experts
A recent publication on arXiv (2608.10605) presents MOSAIC, a novel framework aimed at the simultaneous design of model architecture and systems tailored for the large-scale pretraining of sparse Mixture-of-Experts (MoE) language models. Unlike traditional methods where architecture and system decisions are made independently, MOSAIC treats this co-design as an optimization challenge. It integrates a predictive scaling law with a performance model that assesses Model FLOPs Utilization (MFU), communication expenses, memory usage, and parallel configuration. Focusing on sparse MoE models, the authors analyze how expert quantity and routing sparsity influence loss and efficiency. This research bridges the divide between compute-optimal and cluster-optimal setups, advocating for systems-aware scaling to enhance training efficiency.
Key facts
- Paper title: Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts
- arXiv ID: 2608.10605
- Announce type: cross
- Introduces MOSAIC, a framework for co-designing model architecture and systems
- MOSAIC couples a predictive scaling law with a calibrated performance model
- Performance model estimates MFU, communication cost, memory footprint, and best parallel layout
- Instantiated for sparse Mixture-of-Experts (MoE) language models
- Scaling law fit on sparse MoE models trained on text data, with sparsity factor as a scaling dimension
Entities
Institutions
- arXiv