When an AI model can achieve high scores by paraphrasing or trivially matching evaluation criteria without genuine understanding.