Chain-of-thought reasoning traces appear interpretable but don't reliably encode step importance; this gap matters for process reward models and other techniques that assume reasoning text reflects functional significance.
This paper investigates whether the text of chain-of-thought reasoning steps actually reveals which steps matter for a model's final answer.