Automated benchmark generation validated by domain experts enables continuous evaluation of clinical AI systems, exposing real gaps (like multi-document synthesis) that static benchmarks miss.
Researchers created BRIE, an automatically-generated benchmark for testing how well AI systems retrieve and synthesize information from patient medical records.