You can stop evaluating agents the moment you have statistical proof one is better, rather than running predetermined game counts—this cuts evaluation costs dramatically while keeping your confidence level mathematically sound.
This paper makes agent evaluation cheaper and faster by combining variance reduction (AIVAT) with statistically valid early stopping (Confidence Sequences). Instead of running a fixed number of games, the method stops as soon as there's enough evidence to declare one agent stronger, reducing evaluation costs by up to 74x while maintaining statistical guarantees.