LLM explanations correlate poorly with measured factor importance; operators relying on them to understand or oversee model decisions may be misled, requiring additional verification methods.
This paper tests whether LLM explanations actually match their decision-making by checking if cited factors are truly necessary (changing them changes outputs) or sufficient (keeping them preserves outputs).