ARTFEED — Contemporary Art Intelligence

RL vs SFT: Representational Differences in Math Reasoning Models

ai-technology · 2026-07-30

A recent study published on arXiv (2607.26119) explores the reasons behind the superior performance of large reasoning models trained through reinforcement learning (RL) compared to those fine-tuned via supervised methods (SFT) in mathematical reasoning tasks. By applying linear probes to hidden states across layers, the researchers discovered that RL models demonstrate greater accuracy in predicting the correctness of answers, suggesting their representations are more linearly separable. Additionally, mean ablation studies indicated that RL models create a hierarchical structure, with deeper layers becoming increasingly essential, whereas SFT models show a more uniform distribution of importance. These results imply that RL training fundamentally alters internal representations for enhanced performance.

Key facts

  • Study compares RL and SFT fine-tuned models on math reasoning
  • Linear probes show RL models have higher accuracy in predicting answer correctness
  • RL models develop hierarchical architecture with critical deeper layers
  • SFT models distribute importance uniformly across layers
  • Research provides mechanistic basis for RL's superior performance
  • Published on arXiv with ID 2607.26119
  • Uses mean ablation studies to analyze layer importance
  • Findings demonstrate RL restructures internal representations

Entities

Institutions

  • arXiv

Sources