When evaluating models, adaptively selecting tests based on previous responses can be quadratically more efficient than non-adaptive testing, but this advantage is fundamentally limited and cannot be exponential.
This paper studies how much interaction helps when testing statistical hypotheses. It compares two evaluation strategies: fixing all tests upfront versus adaptively choosing tests based on earlier results. The key finding is that interaction can reduce the number of required tests by a quadratic factor, but not exponentially—despite the apparent complexity of adaptive branching.