ARTFEED — Contemporary Art Intelligence

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

ai-technology · 2026-07-27

Researchers have introduced Re-FORC, a method for adaptive reward prediction that estimates future rewards by analyzing the quantity of thinking tokens. Utilizing a lightweight adapter trained on reasoning models enhances predictions, particularly with extended reasoning and larger models. Re-FORC facilitates the early termination of ineffective reasoning chains, leading to a compute reduction of as much as 26% while preserving accuracy. Additionally, it fine-tunes the selection of model and reasoning length, resulting in a peak accuracy increase of 1.7 percentage points with up to 12% less compute than the largest model alone. Furthermore, adaptive test-time scaling boosts accuracy by 9.9 percentage points compared to confidence-based baselines, allowing for dynamic reasoning controlled by cost-per-token thresholds.

Key facts

  • Re-FORC predicts expected future rewards as a function of future thinking tokens.
  • A lightweight adapter is trained on reasoning models.
  • Early stopping reduces compute by up to 26% with maintained accuracy.
  • Optimized selection achieves 1.7 percentage points higher peak accuracy.
  • Up to 12% less compute needed to match the largest model's accuracy.
  • Adaptive test-time scaling increases accuracy by 9.9 percentage points.
  • Dynamic reasoning with length control via cost-per-token thresholds.
  • Method is proposed in arXiv:2511.02130.

Entities

Institutions

  • arXiv

Sources