REDE: A Framework for Denoising Reasoning Traces to Detect Hallucinations in Large Reasoning Models
A novel framework named REDE has been introduced in a study on arXiv, aimed at enhancing the identification of hallucinations in large reasoning models (LRMs). While LRMs generate extensive reasoning paths, these often include redundant elements that obstruct effective detection of inaccuracies. Traditional methods, such as confidence metrics and basic embedding techniques, struggle to distinguish useful content from irrelevant data. REDE employs final-answer attention to improve representations at each reasoning step, facilitating the identification of noise. The study also highlights two common types of reasoning noise impacting detection and recommends the use of attention mechanisms for better accuracy in truth assessments.
Key facts
- Paper arXiv:2607.22098 introduces REDE framework
- REDE denoises reasoning traces for hallucination detection
- Two forms of reasoning noise identified: irrelevant steps and repetitive steps
- Existing confidence-based scores and naive embedding-based filtering fail to separate noisy from informative steps
- REDE uses final-answer attention as automatic supervision signal
- Framework shapes step-level representation space
- Noisy steps can be reliably identified in refined embeddings
- Research published on arXiv
Entities
Institutions
- arXiv