Language models often cite missing facts when rejecting candidates, but careful testing shows these stated reasons have limited causal influence on their actual choices, raising questions about whether models are genuinely reasoning or post-hoc rationalizing.
This paper tests whether language models' stated reasons for rejecting candidates actually influence their decisions.