Uniform Spectral Growth under Factor-wise Gradient Orthogonalization in Matrix Factorization

Abstract

Spectral gradient descent (SpecGD) orthogonalizes matrix parameter updates and has inspired practical optimizers such as Muon. They often perform well in large language model training, but their dynamics remain poorly understood, especially in factorized parameterizations where the product matrix does not receive orthogonalized updates. We study such dynamics through matrix factorization (MF), where the orthogonalization is applied separately to the factor updates. We analyze spectral gradient flow (SpecGF)—a continuous-time analog of SpecGD—in the low-rank MF setting and prove “equal-rate” dynamics: all singular values grow at equal rates up to small deviations. Consequently, smaller singular values attain their target values earlier than larger ones, contrasting with the largest-first stepwise learning observed in standard gradient flow. Moreover, we prove that SpecGF in our setting converges to global minima from almost all initializations, provided the factor norms remain bounded; with $\ell_2$regularization, we obtain global convergence. Empirically, we observe that LoRA fine-tuning with orthogonalization-based optimizers including Muon exhibit near-uniform growth in the product of LoRA adapters, consistent with the mechanism predicted by our MF analysis.

Publication
Neural Information Processing Systems 2026
Chulhee Yun
Chulhee Yun
Associate Professor

I am an Associate Professor at KAIST AI. I am interested in optimization and machine learning theory.