ARTFEED — Contemporary Art Intelligence

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning

ai-technology · 2026-08-11

A recent paper on arXiv (2608.09324v1) presents CoRE (Consensus Rewards via Equilibrium), a novel approach for reinforcement learning at test time using unlabeled data. Unlike conventional methods that rely on majority voting from N sampled responses—thereby neglecting correct minority responses and treating all majority-matching roll-outs equally—CoRE constructs a graph from these roll-outs. It incorporates edges based on agreement, reasoning similarity, and generation confidence. Dominant sets are identified through replicator dynamics, resulting in a refined pseudo-label, a graded reward for each roll-out, and a cohesiveness gate for each question. CoRE encompasses majority voting as a specific instance. The paper also discusses a block-value analysis that establishes a threshold for achieving consensus to identify correct minority responses amid larger incorrect majorities, with confidence calibration enhancing this threshold. Authored by researchers, this work addresses a key issue in reinforcement learning regarding the lack of ground-truth rewards during testing, aiming to enhance reward estimation and decision-making. Theoretical insights and potential experimental validation are included, although the abstract emphasizes theoretical aspects. This research is significant for the AI and machine learning fields, especially for those focused on reinforcement learning and test-time adaptation.

Key facts

  • Paper ID: arXiv:2608.09324v1
  • Title: CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning
  • Method: CoRE uses graph-based consensus with replicator dynamics
  • CoRE generalizes majority voting
  • Block-value analysis gives threshold for correct minority recovery
  • Confidence calibration lowers threshold multiplicatively
  • Addresses lack of ground-truth reward in test-time RL
  • Published on arXiv

Entities

Institutions

  • arXiv

Sources