AI time horizon benchmarks need better statistical foundations: a 10x increase in human task time doesn't represent equal difficulty gains across all ranges, which matters for fairly comparing AI capabilities.
This paper examines how AI capabilities are measured using 'time horizons'—the human task completion time at which an AI succeeds 50% of the time. The authors show that the standard linear model for this relationship is flawed and propose better statistical methods using splines and item-response theory, revealing that task difficulty doesn't scale uniformly with human time.