You don't need to find the perfect SFT-RL budget split—a wide range of allocations work nearly equally well, and you can identify this range using small models and apply it to large ones, saving computation.
This paper solves a practical problem in LLM training: how to split your annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL). Instead of finding one perfect ratio, the authors identify a 'near-optimal region'—a range of budget splits that all perform nearly as well.