You can predict post-training performance of coding agents by analyzing base models' probability of generating specific successful code actions from recorded trajectories—avoiding expensive full post-training runs.
This paper proposes methods to predict which base language models will perform well after expensive post-training for coding agents, without running the full post-training process.