SFT isn't inherently worse than RL for posttraining—the gap comes from data distribution mismatch. By reshaping data to be more on-policy before training, SFT can generalize better and forget less than strong RL baselines.
This paper shows that supervised finetuning (SFT) can match or beat reinforcement learning for posttraining if you transform the training data first. The authors use an MCMC sampling algorithm to gradually shift off-policy expert demonstrations toward on-policy trajectories that a reference model can actually learn from.