An optimizer's implicit bias toward low-rank solutions depends on whether it respects the gauge symmetry of the loss—a mathematical property independent of the loss value itself.
This paper reveals why different optimizers find different solutions to the same problem: gradient descent implicitly prefers low-rank solutions due to gauge symmetry, while Adam does not. The key difference is whether an optimizer is 'gauge-equivariant'—able to respect the mathematical symmetry of the loss function.