Chulhee Yun

Chulhee Yun

Associate Professor

Kim Jaechul Graduate School of AI

KAIST

My name is Chulhee (I go by Charlie), and I am an Associate Professor at KAIST Kim Jaechul Graduate School of AI (KAIST AI). I also hold a joint affiliation with KAIST Graduate School of AI for Math.

I direct the Optimization & Machine Learning (OptiML) Laboratory at KAIST AI. We study the theoretical principles underlying optimization and deep learning, aiming to bridge the gap between mathematical theory and the practical behavior of modern neural network training. Our work also focuses on developing principled methods and algorithms for modern machine learning.

I received my PhD from the Laboratory for Information and Decision Systems at Massachusetts Institute of Technology, where I was fortunate to study under the joint supervision of Prof. Suvrit Sra and Prof. Ali Jadbabaie. Before MIT, I was a master’s student in Electrical Engineering at Stanford University, where I worked with Prof. John Duchi. I finished my undergraduate program in Electrical Engineering at KAIST.

For prospective students: I look for self-motivated graduate students with strong math and computer science backgrounds. If you are an undergraduate student interested in joining our lab, consider applying for summer/winter KAIST AI Research Internship (KAIRI) programs.

Note for 2027 applicants: I will be on sabbatical at UC Berkeley from January to November 2027. The lab will continue operating during this period, although I may take on fewer new students and interns than usual. Applications are still welcome.

Email: {firstname}.{lastname}@kaist.ac.kr
Phone: +82-2-958-0765
Office: KAIST Seoul Campus Building #9, 9401

Interests
  • Deep Learning Theory
  • Optimization
  • Machine Learning Theory
Education
  • PhD in Elec. Eng. & Comp. Sci., 2016–2021

    Massachusetts Institute of Technology

  • MSc in Electrical Engineering, 2014–2016

    Stanford University

  • BSc in Electrical Engineering, 2007–2014

    KAIST

News

[Aug 2026] I had the pleasure of giving an invited talk at the NUS IMS Workshop on Mathematical Foundations of AI Models. I will also give an invited talk at the upcoming BIRS Workshop on Mathematical Foundations in Deep Learning and Generative AI.
[Jul 2026] I will serve as an organizer of the NeurIPS 2026 Workshop on Optimization for Machine Learning. See you in Sydney!
[May 2026] I was honored to receive the 2026 KSIAM Outstanding Young Investigator Award!
[Jan 2026] Four papers got accepted to ICLR 2026. Kudos to my co-authors!
[Jan 2026] I will give an invited talk at the Workshop on Functional Inference and Machine Intelligence (FIMI 2026) in March 2026.

Publications

Label-Efficient Dataset Pruning via Semi-Supervised Pseudo-Labeling  arXiv
arXiv preprint
Nesterov Acceleration with Operator Decomposition  arXiv
arXiv preprint
AMUSE: Anytime Muon with Stable Gradient Evaluation  arXiv
ICML 2026 Workshop on High-dimensional Learning Dynamics (HiLD)
Uniform Spectral Growth under Factor-wise Muon Orthogonalization in Matrix Factorization and LoRA  arXiv
ICML 2026 Workshop on High-dimensional Learning Dynamics (HiLD)
Transformers Can Learn Multiclass Classification In-Context: Isotropy Governs Generalization 
ICML 2026 Workshop on High-dimensional Learning Dynamics (HiLD)
Understanding Polyak’s Momentum in Deep Learning May Require Rethinking Non-Convex Optimization 
ICML 2026 Workshop on High-dimensional Learning Dynamics (HiLD)
Provably Efficient Regularized Online RLHF with Generalized Bilinear Preferences  arXiv
ICML 2026 Pluralistic Alignment Workshop
Layer Verification Accelerates Speculative Tree Decoding 
ICML 2026 Workshop on Resource-Adaptive Foundation Model Inference (AdaptFM)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime  Paper arXiv
ICLR 2026
NeurIPS 2025 Workshop on Optimization for Machine Learning (OPT 2025)
Implicit Bias and Loss of Plasticity in Matrix Completion: Depth Promotes Low-Rankness  Paper arXiv
ICLR 2026
NeurIPS 2025 Workshop on Dynamics at the Frontiers of Optimization, Sampling, and Games (DynaFront)
ICML 2025 Workshop on High-dimensional Learning Dynamics (HiLD)
The Cost of Robustness: Tighter Bounds on Parameter Complexity for Robust Memorization in ReLU Nets  Paper arXiv
NeurIPS 2025
ICML 2025 Workshop on High-dimensional Learning Dynamics (HiLD)
KAIA Outstanding Paper Award at KAIA Summer Conference 2025 (CKAIA 2025)
From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature Learning  Paper arXiv
NeurIPS 2025
ICML 2025 Workshop on High-dimensional Learning Dynamics (HiLD)
KT Best Paper Award at KAIA Summer Conference 2025 (CKAIA 2025)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training  Paper arXiv
NeurIPS 2025
ICML 2025 Workshop on High-dimensional Learning Dynamics (HiLD)
Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty  Paper arXiv
ICML 2025
ICLR 2025 Workshop on Navigating and Addressing Data Problems for Foundation Models
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification  Paper arXiv
ICLR 2025
Best Paper Award at KAIA Fall Conference 2024 (JKAIA 2024)
Parameter Expanded Stochastic Gradient Markov Chain Monte Carlo  Paper arXiv
ICLR 2025
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count  Paper arXiv
ICLR 2025
Does SGD really happen in tiny subspaces?  Paper arXiv
ICLR 2025
ICML 2024 Workshop on High-dimensional Learning Dynamics 2024: The Emergence of Structure and Reasoning
DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity  Paper arXiv
NeurIPS 2024
ICML 2024 Workshop on Advancing Neural Network Training: Computational Efficiency, Scalability, and Resource Optimization
Provable Benefit of Cutout and CutMix for Feature Learning  Paper arXiv
NeurIPS 2024 (Spotlight)
ICML 2024 Workshop on High-dimensional Learning Dynamics 2024: The Emergence of Structure and Reasoning
KT Best Paper Award at KAIA Summer Conference 2024 (CKAIA 2024)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure  Paper arXiv
NeurIPS 2024
ICML 2024 Workshop on Long-Context Foundation Models
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults  arXiv
ICML 2024 Workshop on High-dimensional Learning Dynamics 2024: The Emergence of Structure and Reasoning
NeurIPS 2023 Workshop on Mathematics of Modern Machine Learning (Oral)
Fundamental Benefit of Alternating Updates in Minimax Optimization  Paper arXiv
ICML 2024 (Spotlight)
ICLR 2024 Workshop on Bridging the Gap Between Practice and Theory in Deep Learning
Linear attention is (maybe) all you need (to understand transformer optimization)  Paper arXiv
ICLR 2024
NeurIPS 2023 Workshop on Mathematics of Modern Machine Learning (Oral)
Fair Streaming Principal Component Analysis: Statistical and Algorithmic Viewpoint  Paper arXiv
NeurIPS 2023
PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning  Paper arXiv
NeurIPS 2023
Practical Sharpness-Aware Minimization Cannot Converge All the Way to Optima  Paper arXiv
NeurIPS 2023 (Spotlight)
KAIA Outstanding Paper Award at KAIA Summer Conference 2023 (CKAIA 2023)
Tighter Lower Bounds for Shuffling SGD: Random Permutations and Beyond  Paper arXiv
ICML 2023 (Oral)
SGDA with shuffling: faster convergence for nonconvex-PŁ minimax optimization  Paper arXiv
ICLR 2023
NAVER Outstanding Theory Paper Award at KAIA-NAVER Joint Fall Conference 2022 (JKAIA 2022)
Minibatch vs Local SGD with Shuffling: Tight Convergence Bounds and Beyond  Paper arXiv
ICLR 2022 (Oral)
Provable Memorization via Deep Neural Networks using Sub-linear Parameters  Paper arXiv
COLT 2021
Presented as part of a contributed talk at DeepMath 2020
A Unifying View on Implicit Bias in Training Linear Neural Networks  Paper arXiv
ICLR 2021
NeurIPS 2020 Workshop on Optimization for Machine Learning (OPT 2020)
Minimum Width for Universal Approximation  Paper arXiv
ICLR 2021 (Spotlight)
Presented as part of a contributed talk at DeepMath 2020
$O(n)$ Connections are Expressive Enough: Universal Approximability of Sparse Transformers  Paper arXiv
NeurIPS 2020
Low-Rank Bottleneck in Multi-head Attention Models  Paper arXiv
ICML 2020
Are Transformers universal approximators of sequence-to-sequence functions?  Paper arXiv
ICLR 2020
NeurIPS 2019 Workshop on Machine Learning with Guarantees
Honorable Mention at NYAS Machine Learning Symposium 2020 Poster Awards
Are deep ResNets provably better than linear predictors?  Paper arXiv
NeurIPS 2019
Small ReLU networks are powerful memorizers: a tight analysis of memorization capacity  Paper arXiv
NeurIPS 2019 (Spotlight)
Minimax Bounds on Stochastic Batched Convex Optimization  Paper
COLT 2018
Global optimality conditions for deep neural networks  Paper arXiv
ICLR 2018
NIPS 2017 Workshop on Deep Learning: Bridging Theory and Practice
Face detection using Local Hybrid Patterns  Paper
ICASSP 2015

Teaching

AI.50500 Optimization for AI (F2025, F2026)
AI.61600 Deep Learning Theory (S/F2022, S/F2023, F2024, S2026)
AI.70900 Advanced Deep Learning Theory (S2024, S2025)

Research Group

I direct the Optimization & Machine Learning (OptiML) Laboratory at KAIST. I ambitiously pronounce it as the “Optimal Lab”—although my students may disagree!

Postdocs/InnoCORE Fellows
  • Yoonsoo Nam, PhD
PhD and MS/PhD Students (all students are in KAIST AI)
Master’s Students (all students are in KAIST AI)
Undergraduate Students/KAIRI Interns
  • Huiwone Kim (SNU ECE)
  • Jihwan Kim (SNU Math/CS)
  • Yelynn Suh (KAIST EE)
Former Graduate Students/Notable Former Interns

Service

Conference/Workshop Organizing Committee
  • Organizer, NeurIPS 2026 Workshop on Optimization for Machine Learning
  • Social Chair, ICML 2026
Conference Area Chair
  • ICLR 2025–2027
  • NeurIPS 2023–2026 (Selected as a Notable AC for NeurIPS 2023)
  • ICML 2026
Conference/Workshop Reviewer
  • ICML 2019–2025 (Selected as a Top Reviewer at ICML 2025)
  • ICLR 2019–2024
  • COLT 2020–2024
  • NeurIPS 2018–2020, 2022
  • AISTATS 2019
  • CDC 2018
  • ICML 2026 Workshop on High-dimensional Learning Dynamics
  • ICML 2025 Workshop on High-dimensional Learning Dynamics
  • ICLR 2025 Workshop on Will Synthetic Data Finally Solve the Data Access Problem?
  • ICLR 2024 Workshop on Privacy Regulation and Protection in Machine Learning
Journal Reviewer
  • Journal of Machine Learning Research
  • SIAM Journal on Mathematics of Data Science
  • Annals of Statistics
  • IEEE Transactions on Neural Networks and Learning Systems
  • IEEE Transactions on Information Theory
  • Mathematical Programming
  • Neural Networks
  • Stochastic Systems
  • Artificial Intelligence
  • Information and Inference: A Journal of the IMA