Understanding Polyak’s Momentum in Deep Learning May Require Rethinking Non-Convex Optimization

Abstract

Polyak’s heavy-ball momentum is widely used in deep learning, where it often accelerates training in practice. However, standard smooth non-convex optimization theory, which typically measures convergence by averaged or best-iterate gradient norms, offers only a limited explanation of this advantage. We revisit this gap through worst-case lower bounds. For SGD with heavy-ball momentum (SHB), SignGD with momentum (Signum), and Muon, we show that the lower bounds on averaged gradient norms considered for these methods can exceed the upper bounds for their non-momentum counterparts, even with their respective optimal constant step sizes. These comparisons cover commonly used ranges of the momentum parameter $\beta$, become unfavorable to momentum as $\beta$ increases, and diverge in deterministic settings as $\beta \rightarrow 1$. For GD with heavy-ball momentum (HB), we further show that the same separation persists under the best-iterate squared gradient norm. These results indicate that the standard framework can lead to comparisons opposite to the empirical behavior of momentum in deep learning, motivating refinements involving convergence measures, structure, or stochasticity.

Publication
ICML 2026 Workshop on High-dimensional Learning Dynamics (HiLD)
Chulhee Yun
Chulhee Yun
Associate Professor

I am an Associate Professor at KAIST AI. I am interested in optimization and machine learning theory.