Gradient clipping in decentralized SGD provably converges at optimal rates under heavy-tailed noise with linear speedup across agents—outperforming normalization-based approaches that require additional techniques like momentum.
This paper proves that decentralized SGD with gradient clipping achieves optimal convergence rates under heavy-tailed noise, a common problem in modern machine learning. The key insight is that clipping preserves gradient magnitude information better than normalization, enabling faster convergence and linear speedup across multiple agents in a decentralized network.