Scaling laws predict transformer performance on collider jets
A new study demonstrates that scaling laws can forecast the performance of transformer models in particle physics before full training. Researchers fit a joint model-and-data scaling law on small models spanning three orders of magnitude of training compute, then predicted the loss of larger models trained with over 100 times more compute to within one percent. The study connects pretraining loss to downstream physics performance: on two standard tagging benchmarks, lower pretraining loss yields lower fine-tuning loss and higher background rejection. This enables compute budgets to be translated into expected performance for this model family and tasks.
Key facts
- Scaling laws predict transformer performance on collider jets
- Joint model-and-data scaling law fit on small models
- Spans three orders of magnitude of training compute
- Predicts loss of larger models with over 100x more compute to within 1%
- Lower pretraining loss yields lower fine-tuning loss
- Higher background rejection after fine-tuning
- Two standard tagging benchmarks used
- Compute budget can be translated into expected performance
Entities
—