Compression Scores Fail to Decide: Information Boundaries for Group-Robust LLM Pruning
A recent study on arXiv (2608.02940) has found that using a reproducible compression statistic might wrongly flag candidates for pruning large language models (LLMs). Even though a dense pruning score showed a reliability of 0.906 and hinted at a 16.1% improvement, the final results demonstrated declines of 6.0% and 7.7% compared to two control groups. The researchers attribute these inconsistencies to the ways different statistics define information interfaces. According to a conic law, the pooling cost for positive linear fixed-candidate damage is exact, incorporating various second moments. Their findings indicate a strong Spearman correlation of 0.9239 for grouped diagonal damages, while finer orders lack robustness. Lastly, a coarser depth allocation can lessen worst-group perplexity by 12.6–20.9% across three LLMs, highlighting the need for better methods in pruning decisions.
Key facts
- arXiv:2608.02940
- Dense pruning score with 0.906 split-half reliability predicted a 16.1% gain
- Selected endpoint was 6.0% and 7.7% worse than two controls
- Conic law gives exact pooling price for positive linear fixed-candidate damage
- Group-resolved diagonal recovers broad damage order (Spearman 0.9239)
- Coarse depth allocation cuts worst-group perplexity inflation by 12.6–20.9%
- Three dense LLMs used in experiments
- Model-specific complete-mask endpoint selection discussed
Entities
Institutions
- arXiv