SG-TULA provides a theoretically-grounded alternative to standard optimizers for non-smooth, non-convex problems with explicit convergence guarantees—useful for understanding and improving LLM pretraining when standard assumptions break down.
This paper introduces SG-TULA, a sampling algorithm for training machine learning models when the loss landscape is non-smooth, non-convex, and has steep gradients. Unlike existing methods, it works directly with subgradients without expensive smoothing, uses taming techniques for stability, and comes with theoretical guarantees on convergence speed.