Beyond Rate Optimality: How Polyak’s Momentum Shapes Non-Convex Optimization

Abstract

Polyak’s heavy-ball momentum is widely used in deep learning, where it often accelerates training in practice, but this advantage is not consistently reflected in standard smooth non-convex optimization theory. We ask whether the unfavorable comparisons in existing convergence guarantees can be solely attributed to their suboptimal analyses. For SGD with heavy-ball momentum (SHB), SignGD with momentum (Signum), and Muon, we derive worst-case lower bounds on averaged gradient norms that can exceed upper bounds for their non-momentum counterparts, even when the constant step size is optimized separately for each method. These comparisons cover commonly used ranges of the momentum parameter $\beta$, become increasingly unfavorable as $\beta$ grows, and diverge in deterministic settings as $\beta \to 1$. For deterministic HB, we further show an analogous separation under the best-iterate squared gradient norm. Thus, the unfavorable comparisons in existing guarantees cannot in general be dismissed as artifacts of loose upper bounds. We discuss additional features that may determine when momentum is beneficial, including convergence criteria, structural assumptions, and stochasticity.

Publication
NeurIPS 2026 Workshop on Optimization for Machine Learning (OPT 2026)
Chulhee Yun
Chulhee Yun
Associate Professor

I am an Associate Professor at KAIST AI. I am interested in optimization and machine learning theory.