LaPrune: Exact-Budget Differentiable Sparsity for Million-Scale Models
There’s a new study out about LaPrune, a unique layer that helps train sparse models while sticking to a specific budget and managing the normalized second moment. Unlike traditional methods that can disrupt gradients, LaPrune uses a LapSum barrier for selection mass and a normalized second-moment constraint. This allows for a transition from evenly distributing mass to strict top-k selections within the budget. The technique guarantees accuracy in budgeting and provides strong theoretical predictions, including insights on saturated fractions and a solid worst-case scenario for near-zero fractions. You can find this research on arXiv with ID 2608.04057 in the Computer Science > Machine Learning section, and it’s particularly relevant for AI and machine learning applications.
Key facts
- LaPrune is a differentiable layer for sparse model training with exact budget control.
- It controls the normalized second moment while preserving selected mass.
- Uses a LapSum barrier to preserve selection mass.
- A normalized second-moment constraint moves the mask from dense equal-mass allocation to hard top-k.
- Provides a population prediction of the saturated fraction.
- Derives a near-binary limiting law.
- Offers a tight worst-case guarantee on the near-zero fraction.
- The normalized hardness parameter is invariant to score scale.
- Paper available on arXiv with identifier 2608.04057.
- Categorized under Computer Science > Machine Learning.
Entities
Institutions
- arXiv