Deceptive outputs from language models don't necessarily mean the model has deceptive intentions or mechanisms—you need causal evidence to distinguish between what a model does and why it does it.
This paper examines whether language models that produce deceptive outputs actually have deceptive mechanisms inside them. The authors create a causal framework to distinguish between behavior that looks deceptive and mechanisms that are genuinely deceptive, then test these distinctions through controlled experiments with language models.