LZ Penalty Rids Autoregressive Models of Degenerate Repetition Without Capability Loss
A recent paper published on arXiv presents the LZ penalty, an innovative information-theoretic approach aimed at removing degenerate repetitions in autoregressive language models while preserving performance. This method, inspired by the LZ77 compression algorithm, enables top-tier open-source reasoning models to function with greedy decoding without encountering degenerate repetition, contrasting with conventional methods that may result in rates of up to 4%. The paper, categorized under Computer Science > Machine Learning with the identifier 2504.20131, bridges the gap between compression theory and text generation, focusing on eliminating repetitive patterns deemed redundant. This technique proves effective across leading models and underscores the wider relevance of information theory in natural language processing, especially for deterministic outputs in automated reasoning and coding tasks.
Key facts
- The LZ penalty is introduced as a repetition penalty for autoregressive language models.
- It is based on codelengths in the LZ77 universal lossless compression algorithm.
- The penalty interprets decoding as sampling from the residual distribution after removing highly compressible information.
- It enables state-of-the-art open-source reasoning models to use greedy decoding without degenerate repetition.
- The method does not cause loss of capability.
- Industry-standard frequency and repetition penalties incur degenerate repetition rates of up to 4%.
- The paper is listed under Computer Science > Machine Learning.
- The preprint is available on arXiv with identifier 2504.20131.
Entities
Institutions
- arXiv
- arXivLabs
- Semantic Scholar