Layer Pruning Fails to Preserve Generative Reasoning in LLMs
A recent arXiv paper (2602.01997v4) investigates the impact of layer pruning on large language models (LLMs), revealing that while classification benchmarks recover well, generative reasoning tasks suffer. The study, conducted under realistic post-training constraints with a single 80GB GPU and no pretraining-scale data, shows that a simple supervised finetuning strategy with self-generated responses recovers up to 90% of baseline performance on classification, but generative reasoning recovery remains limited. The authors demonstrate that pruning leads to a loss of key algorithmic capabilities, such as arithmetic computation and balanced parenthesis generation, beyond surface-level text degradation. This gap persists even under task-aligned recovery, indicating fundamental limitations of layer pruning for generative reasoning. The paper is available on arXiv under the identifier 2602.01997.
Key facts
- Layer pruning can compress LLMs while retaining performance on classification benchmarks.
- Generative reasoning tasks like GSM8K and HumanEval+ show substantially weaker recovery.
- Pruning leads to loss of arithmetic computation and balanced parenthesis generation.
- Recovery strategy uses supervised finetuning with self-generated responses.
- Recovery achieves up to 90% of baseline on classification but limited on generative reasoning.
- Experiments conducted with a single 80GB GPU and without pretraining-scale data.
- The gap persists even under task-aligned recovery.
- Paper available on arXiv (2602.01997v4).
Entities
Institutions
- arXiv