Survey on Low-Precision Training of Large Language Models
A comprehensive survey on low-precision training methods for large language models (LLMs) has been released on arXiv. The paper, identified as arXiv:2505.01043v2, addresses the challenge of substantial hardware resources required for LLM training by reviewing low-precision techniques that improve efficiency. The authors categorize existing methods into three primary groups based on the numerical format used for weights, activations, and gradients. The survey aims to provide a unified overview of a fragmented research landscape, highlighting the diversity of numerical representations and the resulting difficulties for researchers. The work covers methods, challenges, and opportunities in low-precision training, offering a systematic organization of approaches to facilitate further advancements.
Key facts
- arXiv paper ID: 2505.01043v2
- Announce type: replace-cross
- Focuses on low-precision training of large language models
- Categorizes methods into three groups based on numerical format
- Addresses hardware resource barriers for LLM training
- Provides a comprehensive review of existing methods
- Highlights fragmented landscape in low-precision training research
- Aims to offer a unified overview for researchers
Entities
Institutions
- arXiv