When evaluating AI systems across domains with limited labels, borrowing information across domains through prediction-powered smoothing gives more accurate performance estimates than evaluating each domain independently.
This paper addresses how to accurately evaluate AI systems across different domains (like task types or conversation types) when you can only label a small sample. It proposes prediction-powered smoothing, which combines AI predictions with limited labels to get better estimates for each domain, and a validation method to choose between different estimation approaches.