Synthetic clinical benchmarks need explicit realism optimization separate from utility validation—passing utility checks alone doesn't guarantee a benchmark reflects real-world data patterns, which matters for training reliable healthcare AI agents.
This paper addresses a critical gap in synthetic clinical benchmarks: they can pass utility checks while remaining structurally unrealistic. The authors develop methods to improve benchmark realism (measured by data missingness patterns, actionability, and population alignment) while maintaining the utility thresholds required for downstream AI systems.