ARTFEED — Contemporary Art Intelligence

ReCo: Reward-Coordinated Compression for Efficient LRM Reasoning

ai-technology · 2026-08-06

A new research paper on arXiv (2608.04771) introduces ReCo (Reward-Coordinated Compression), a framework for reducing inference costs in Large Reasoning Models (LRMs) by compressing KV-cache. The paper observes that reasoning states vary in tolerance to context loss along the trajectory, and process reward can track this: deleting tokens at high-reward steps preserves accuracy better than random deletion. Additionally, compression on the generation side is not free, as a smaller cache leads to more generated tokens, partially offsetting savings. ReCo coordinates both compression and generation under a single process reward. The paper is authored by researchers (names not provided in the abstract) and was announced as new on arXiv. The work addresses the problem of overthinking in LRMs, which inflates inference costs due to lengthy chain-of-thought reasoning. The proposed method aims to improve efficiency without sacrificing accuracy.

Key facts

  • Paper on arXiv: 2608.04771
  • Introduces ReCo (Reward-Coordinated Compression)
  • Targets KV-cache compression for Large Reasoning Models
  • Observations: reasoning state tolerance varies, process reward tracks it
  • Compression leads to more generated tokens, partially canceling savings
  • Coordinates compression and generation under a single process reward
  • Addresses overthinking in LRMs
  • Aims to reduce inference cost while preserving accuracy

Entities

Institutions

  • arXiv

Sources