New Scaling Laws for Compute- and Data-Optimal Pretraining
A new arXiv paper proposes Compute-Data (CD) scaling laws to address the growing imbalance between compute and high-quality pretraining data. Classical scaling laws assume unlimited fresh data, but as compute outpaces data availability, CD scaling bridges compute-optimal (data scales with compute) and data-optimal (fixed corpus, unlimited compute) regimes. The framework introduces a token-effectiveness function η, quantifying the value of derived tokens from multi-epoch repetition or paraphrasing relative to fresh tokens. Experiments on the Dolma-3 corpus with models from 14M to 600M parameters show token effectiveness is far from ideal.
Key facts
- Paper title: Bridging Compute- and Data-Optimal Pretraining
- arXiv ID: 2607.25271
- Classical scaling laws assume unbounded fresh data
- CD scaling unifies compute-optimal and data-optimal scaling
- Introduces token-effectiveness function η
- η measures value of derived tokens vs fresh tokens
- Tested on Dolma-3 corpus
- Model sizes: 14M to 600M parameters
- Data-expansion strategies: multi-epoch repetition and paraphrasing
- Token effectiveness far from perfect
Entities
Institutions
- arXiv