arXiv Study: Data Selection Effects Depend on Model Capacity
A recent study published on arXiv (2608.13721) examines how the success of likelihood-based data selection in supervised fine-tuning is influenced by the model's capacity and the duration of training. Researchers, utilizing models with parameters ranging from 1.5B to 8B and guidance from more advanced teacher models, identify a 'Fast-Fit / Slow-Gain' trend that depends on capacity. While high-likelihood data leads to quicker and more consistent initial enhancements, the long-term advantages are contingent upon the size of the model and the length of training. The research concentrates on mathematical reasoning tasks and includes controlled experiments to validate its conclusions. The paper can be accessed at https://arxiv.org/abs/2608.13721.
Key facts
- Paper arXiv:2608.13721 examines likelihood-based data selection in reasoning supervised fine-tuning.
- Effectiveness depends on model capacity and training duration.
- Students range from 1.5B to 8B parameters.
- Supervision generated by stronger teacher models.
- Observed capacity-dependent 'Fast-Fit / Slow-Gain' pattern.
- High-likelihood data provides faster and more stable early improvements.
- Study focuses on mathematical reasoning.
- Controlled experiments conducted.
- Paper available on arXiv.
Entities
Institutions
- arXiv