Evaluation metrics for AI radiology report generation are sensitive to reporting style choices, not just clinical accuracy—picking the right reference reports matters as much as the model itself.
This paper reveals that how radiologists write reports—their choice of words, detail level, and formatting—significantly affects how AI-generated radiology reports are evaluated.