When an LLM's stated reasoning or extracted language features don't reflect how the model actually computes internally.