Even when an optimizer finds a correct solution to a task, different optimizers can cause the solution to collapse post-training due to mismatched step-size dynamics at the representation-readout interface—a failure mode invisible to standard loss curves.
This paper investigates why Muon-optimized transformers solve modular arithmetic tasks during training but then catastrophically lose generalization afterward. The authors identify the failure occurs at the interface between learned representations and the output layer, where different optimizer dynamics cause instability.