TAM: Thought-Aware KV Cache Compaction for Reasoning Models
A novel approach known as Thought-Aware Attention Matching (TAM) tackles the memory limitations in reasoning language models by leveraging the hierarchical nature of chain-of-thought (CoT) sequences. In contrast to current compaction techniques that view reasoning paths as flat token sequences with uniform compression, TAM breaks down the trajectory into reasoning blocks through thought segmentation. It allocates compression budgets adaptively, considering the significance and size of each segment, while safeguarding crucial high-attention tokens. The authors demonstrate that the allocation strategy is optimal within a convex error framework and that cumulative errors from sequential compaction remain limited. Experiments were carried out on the AIME 2024 and MATH datasets, though specific findings are not included in the abstract. The paper can be found on arXiv with the identifier 2608.12331.
Key facts
- TAM is a novel KV cache compaction method for reasoning language models.
- It uses thought segmentation to decompose CoT trajectories into reasoning blocks.
- Adaptive budget allocation assigns compression budget based on segment importance and size.
- Pivotal token protection preserves high-attention reasoning anchors.
- The allocation rule is proven optimal under a convex error model.
- Cumulative error under sequential compaction remains bounded.
- Experiments were performed on AIME 2024 and MATH datasets.
- The paper is available on arXiv (2608.12331).
Entities
Institutions
- arXiv