Test-time scaling isn't one thing: different inference strategies have different statistical structures and failure modes. Comparing them requires protocol-matched reporting of compute, uncertainty, and the exact inference method used—not just a single 'budget' number.
This paper systematizes how LLMs can solve harder problems by spending more compute at inference time. It clarifies that different 'test-time scaling' approaches (like trying multiple solutions or searching through partial answers) work fundamentally differently, making direct comparisons misleading.