Steering Vectors for Chain-of-Thought Faithfulness: Generalization Study
A new arXiv preprint (2607.29062) investigates how activation steering can improve the faithfulness of chain-of-thought (CoT) reasoning in large language models. The authors extend prior work by examining the generalization of steering vectors across different cue types, datasets, and construction methods. They test three models: Gemma-3 4B, Qwen-3.5 9B, and Gemma-3 12B, in a cued question-answering setting. The study addresses the problem where models fail to verbalize instrumental reasoning steps, especially when prompted with misleading cues. The findings indicate that steering reliably increases faithfulness, though the degree of generalization varies. This research contributes to AI safety by enhancing the transparency of model reasoning.
Key facts
- arXiv preprint 2607.29062
- Study on steering vectors for chain-of-thought faithfulness
- Models tested: Gemma-3 4B, Qwen-3.5 9B, Gemma-3 12B
- Cued question-answering setting
- Generalization across cue types, datasets, and steering vector construction methods
- Steering reliably increases faithfulness
- Addresses unfaithful CoT where models omit instrumental reasoning steps
- Prior work showed activation steering improves faithfulness
Entities
Institutions
- arXiv