Publications

Nesterov Acceleration with Operator Decomposition. arXiv preprint, 2026.
Stronger Lower Bounds for (Non-)Anytime Acceleration of Gradient Descent. NeurIPS 2026 Workshop on Optimization for Machine Learning (OPT 2026), 2026.
From Geometry to Generalization: Why Row Normalization Can Beat Adam and Muon. NeurIPS 2026 Workshop on Optimization for Machine Learning (OPT 2026), 2026.
Beyond Loss Gaps: Stability and Geometry in Dynamic Sample Reweighting. NeurIPS 2026 Workshop on Optimization for Machine Learning (OPT 2026), 2026.
The Phase Transition in Random Reshuffling: Tight Rates under Strong Convexity. NeurIPS 2026 Workshop on Optimization for Machine Learning (OPT 2026), 2026.
Machine Unlearning with a Destination: Tracking a Minimizer Path to the Retain-Only Objective. NeurIPS 2026 Workshop on Optimization for Machine Learning (OPT 2026), 2026.
Beyond Rate Optimality: How Polyak’s Momentum Shapes Non-Convex Optimization. NeurIPS 2026 Workshop on Optimization for Machine Learning (OPT 2026), 2026.
In-Context Learning of Hidden Markov Models by Loop Transformers. NeurIPS 2026 Workshop on AI for Stochastic Dynamics (STODY), 2026.
AMUSE: Anytime Muon with Stable Gradient Evaluation. NeurIPS 2026 (Spotlight), 2026.
Uniform Spectral Growth under Factor-wise Gradient Orthogonalization in Matrix Factorization. NeurIPS 2026, 2026.
Label-Efficient Dataset Pruning via Semi-Supervised Pseudo-Labeling. NeurIPS 2026, 2026.
Looking Through the Mirror: Minimax-Optimal Regularized Regrets in Online Learning and Bandits. NeurIPS 2026, 2026.
Provably Efficient Regularized Online RLHF with Generalized Bilinear Preferences. NeurIPS 2026, 2026.
Layer Verification Accelerates Speculative Tree Decoding. ICML 2026 Workshop on Resource-Adaptive Foundation Model Inference (AdaptFM), 2026.
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?. ICLR 2026, 2026.
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training. NeurIPS 2025, 2025.
Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty. ICML 2025, 2025.
Parameter Expanded Stochastic Gradient Markov Chain Monte Carlo. ICLR 2025, 2025.
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count. ICLR 2025, 2025.
Does SGD really happen in tiny subspaces?. ICLR 2025, 2025.
DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity. NeurIPS 2024, 2024.
Provable Benefit of Cutout and CutMix for Feature Learning. NeurIPS 2024 (Spotlight), 2024.
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure. NeurIPS 2024, 2024.
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults. ICML 2024 Workshop on High-dimensional Learning Dynamics 2024: The Emergence of Structure and Reasoning, 2024.
Fundamental Benefit of Alternating Updates in Minimax Optimization. ICML 2024 (Spotlight), 2024.
Linear attention is (maybe) all you need (to understand transformer optimization). ICLR 2024, 2023.
Fair Streaming Principal Component Analysis: Statistical and Algorithmic Viewpoint. NeurIPS 2023, 2023.
PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning. NeurIPS 2023, 2023.
Practical Sharpness-Aware Minimization Cannot Converge All the Way to Optima. NeurIPS 2023 (Spotlight), 2023.
Tighter Lower Bounds for Shuffling SGD: Random Permutations and Beyond. ICML 2023 (Oral), 2023.
On the Training Instability of Shuffling SGD with Batch Normalization. ICML 2023, 2023.
Minibatch vs Local SGD with Shuffling: Tight Convergence Bounds and Beyond. ICLR 2022 (Oral), 2022.
Provable Memorization via Deep Neural Networks using Sub-linear Parameters. COLT 2021, 2021.
A Unifying View on Implicit Bias in Training Linear Neural Networks. ICLR 2021, 2021.
Minimum Width for Universal Approximation. ICLR 2021 (Spotlight), 2021.
SGD with shuffling: optimal rates without component convexity and large epoch requirements. NeurIPS 2020 (Spotlight), 2020.
$O(n)$ Connections are Expressive Enough: Universal Approximability of Sparse Transformers. NeurIPS 2020, 2020.
Low-Rank Bottleneck in Multi-head Attention Models. ICML 2020, 2020.
Are Transformers universal approximators of sequence-to-sequence functions?. ICLR 2020, 2020.
Are deep ResNets provably better than linear predictors?. NeurIPS 2019, 2019.
Small ReLU networks are powerful memorizers: a tight analysis of memorization capacity. NeurIPS 2019 (Spotlight), 2019.
Efficiently testing local optimality and escaping saddles for ReLU networks. ICLR 2019, 2019.
Minimax Bounds on Stochastic Batched Convex Optimization. COLT 2018, 2018.
Global optimality conditions for deep neural networks. ICLR 2018, 2018.
Face detection using Local Hybrid Patterns. ICASSP 2015, 2015.