ARTFEED — Contemporary Art Intelligence

RAFS: A New Score to Detect Silent Reasoning Failures in LLMs

ai-technology · 2026-07-30

A recent paper presents the Reasoning Answer Faithfulness Score (RAFS), which serves as a reference-free diagnostic tool for identifying silent reasoning errors in large language models at the instance level. RAFS tackles the inconsistency between reasoning and answers, where a model might arrive at a correct answer through an invalid reasoning path or a valid calculation that is marred by a transcription mistake. This score integrates factors such as step validity, entailment from reasoning to answer, sensitivity to counterfactuals, consensus on answers, and stability in conditional reasoning. It focuses on the agreement at the transcript level instead of assessing the model's internal computations or factual accuracy beyond the specific mathematical context. The paper can be found on arXiv with ID 2607.26102.

Key facts

  • RAFS is a reference-free score for detecting reasoning failures in LLMs.
  • It addresses the reasoning-answer consistency gap.
  • RAFS combines step validity, entailment, counterfactual sensitivity, answer consensus, and conditional reasoning stability.
  • It evaluates transcript-level agreement, not private computation.
  • The paper is on arXiv with ID 2607.26102.
  • It is a framework paper introducing the diagnostic.
  • RAFS is instance-level and targeted.
  • It does not assess factual correctness outside the mathematical setting.

Entities

Institutions

  • arXiv

Sources