Generative reward models work better in RL when you extract rewards from their ranking capabilities rather than forcing them to produce scalar scores—use pairwise or reference-based comparisons instead.
This paper shows how to use generative reward models (which rank responses) effectively in reinforcement learning for language models. The key insight is that generative models naturally compare responses rather than score them individually.