Evaluating AI outputs requires measuring both correctness (precision) and completeness—whether responses include all relevant facts. This benchmark reveals that current models often miss important information even when what they do say is accurate.
This paper introduces GAMUT, a benchmark for evaluating whether long-form AI-generated text contains all necessary information (factual completeness), not just whether claims are correct. The authors propose a two-level rubric framework that captures content organization and importance, then converts it into machine-gradable checklists.