Muon optimizer can substantially improve agentic RL performance, but its effectiveness is tightly coupled to the choice of advantage estimator and learning rate—practitioners should tune these jointly rather than treating them independently.
This paper investigates when the Muon optimizer helps in reinforcement learning for agentic tasks. Testing on ALFWorld with a small language model, the authors find that Muon significantly outperforms AdamW in sparse-reward settings—improving success rates by 88% in some configurations.