Speech-based disease detection models may not be learning genuine disease characteristics but rather dataset-specific patterns, raising serious concerns about their reliability for real-world clinical use.
This paper investigates whether speech models trained to detect Parkinson's disease actually learn disease-specific patterns or just exploit dataset quirks.