Counterfactual Simulation Training Enhances Chain-of-Thought Faithfulness
Counterfactual Simulation Training (CST) is a novel approach designed to enhance the reliability of Chain-of-Thought (CoT) reasoning in large language models (LLMs). This technique, detailed in an arXiv paper (2602.20710), incentivizes CoTs that help a simulator effectively forecast a model's outputs in response to counterfactual inputs. CST is utilized in two contexts: monitoring CoT with cue-based counterfactuals to identify dependence on misleading features, reward manipulation, or sycophantic behavior, and conducting counterfactual simulations using generic model-based counterfactuals to promote more accurate and adaptable reasoning. Tests on models with up to 235B parameters revealed a significant 35-point increase in monitoring accuracy for cue-based counterfactuals. The paper tackles recognized issues regarding CoT faithfulness that hinder the understanding of reasoning processes.
Key facts
- CST is a training method to improve CoT faithfulness.
- It rewards CoTs that enable accurate prediction of outputs over counterfactual inputs.
- Applied in two settings: CoT monitoring with cue-based counterfactuals and counterfactual simulation over generic model-based counterfactuals.
- Detects reliance on spurious features, reward hacking, and sycophancy.
- Experiments with models up to 235B parameters.
- Improves monitor accuracy on cue-based counterfactuals by 35 accuracy points.
- Paper available on arXiv with ID 2602.20710.
- Addresses problems with CoT faithfulness in LLMs.
Entities
Institutions
- arXiv