When generating synthetic medical images, use domain-specific evaluation metrics rather than generic ones—and prioritize generating diverse data over pixel-perfect realism for downstream medical AI tasks.
This paper evaluates how well synthetic histopathology images generated by diffusion models work for medical AI tasks. The authors show that standard image quality metrics (FID, IS) designed for natural images fail for medical images, and propose using pathology-specific metrics instead.