Using spectral norm geometry in SAM's perturbation step combined with Muon optimization leads to better generalization on vision models—a practical improvement for training robust neural networks.
This paper improves Sharpness-Aware Minimization (SAM), a training technique that helps models generalize better, by using matrix-aware geometry. The authors propose using spectral norm-based perturbations for hidden-layer weights and combine this with the Muon optimizer, achieving better results on ImageNet with ViT and ResNet models.