Adaptive token-level iteration selection during inference can significantly improve test-time scaling efficiency—TaH2 achieves 53% better accuracy-per-compute gains than fixed looping by intelligently deciding which tokens deserve extra processing passes.
This paper improves how AI models use extra computation time during inference by introducing TaH2, which selectively applies multiple processing passes to tokens that benefit most from them. Unlike standard looped transformers that process every token multiple times, TaH2 learns which tokens need extra iterations, achieving better accuracy gains per unit of compute on math reasoning tasks.