Automated TTS evaluators need to move beyond generic 'naturalness' scores and evaluate specific linguistic dimensions of speech quality, as current tools miss many errors that human listeners easily detect.
This paper reveals that current automated Text-to-Speech evaluation tools (MOS predictors and Audio-LLM judges) fail to capture the full range of speech quality issues that humans perceive.