Assessing AI system performance separately across different subgroups or domains rather than reporting a single overall metric.