ARTFEED — Contemporary Art Intelligence

MARS: Adaptive Data Augmentation Framework for Reward Modeling in RLHF

ai-technology · 2026-08-07

Researchers have introduced a new framework called MARS, which stands for Margin and Semantic-Aware Data Augmentation for Reward Modeling, aimed at boosting reward modeling in scenarios with limited data. Reward modeling is crucial in techniques like reinforcement learning from human feedback (RLHF) and AI feedback (RLAIF), but it often struggles with insufficient and inconsistent human preference data. MARS tackles this by emphasizing augmentation for low-margin preference pairs and enhancing semantic distinctions to effectively separate accepted from rejected responses before generating synthetic preference samples. Evaluated across three datasets and two reward-modeling frameworks, MARS showed better performance in RewardBench and alignment win rates compared to traditional methods. The full details are available on arXiv under the identifier 2602.17658 in the Machine Learning section.

Key facts

  • MARS stands for Margin and Semantic-Aware Data Augmentation for Reward Modeling.
  • It is an adaptive augmentation framework for controlled low-resource reward modeling.
  • MARS allocates more augmentation to low-margin preference pairs.
  • It uses semantic-distance-based refinement to improve chosen-rejected contrast.
  • Evaluated on three preference datasets and two reward-model backbones.
  • Improves average RewardBench performance and alignment win rates.
  • Compared against uniform augmentation, WoN, and AdaBoost-style baselines.
  • Gains are not solely explained by semantic refinement or GPT-4.1 judge coupling.

Entities

Institutions

  • arXiv

Sources