Quantizing optimizer states in preconditioner space (where learning rates are computed) rather than state space reduces training loss gaps by up to 70%, making 4-bit AdamW practical for large-scale pretraining without sacrificing convergence.
This paper improves 4-bit quantization of AdamW optimizer states by rounding in preconditioner space instead of state space. The authors show that quantization errors in the second moment (used to scale learning rates) cause larger problems than errors in the first moment, and propose two methods—ZIP-SR and ZE-EDEN—that better preserve the adaptive learning rates.