When auditing AI models for medical text classification, use multiple explanation methods together and check their agreement—single methods can be misleading, especially when the model is uncertain about its prediction.
This paper evaluates how well different explanation methods agree when analyzing a medical text classifier (DeBERTa-v3). Using five different explanation techniques on medical abstracts, researchers found that explanations are most reliable when the model is confident, but diverge significantly when the model is uncertain.