8 papers
Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware Minimization
Chaewon Moon, Dongkuk Si, Chulhee Yun
We study the implicit bias of Sharpness-Aware Minimization (SAM) when training -layer linear diagonal networks on linearly separable binary classification. For linear models ($L…
Implicit Bias and Loss of Plasticity in Matrix Completion: Depth Promotes Low-Rankness
Baekrok Shin, Chulhee Yun
We study matrix completion via deep matrix factorization (a.k.a. deep linear neural networks) as a simplified testbed to examine how network depth influences training dynamics. Des…
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
Jihwan Kim, Dogyoon Song, Chulhee Yun
We study scaling laws of signSGD under a power-law random features (PLRF) model that accounts for both feature and target decay. We analyze the population risk of a linear model tr…
Uniform Spectral Growth and Convergence of Muon in LoRA-Style Matrix Factorization
Changmin Kang, Jihun Yun, Baekrok Shin +2
Spectral gradient descent (SpecGD) orthogonalizes the matrix parameter updates and has inspired practical optimizers such as Muon. They often perform well in large language model (…
From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature Learning
Junsoo Oh, Jerry Song, Chulhee Yun
Weak-to-strong generalization refers to the phenomenon where a stronger model trained under supervision from a weaker one can outperform its teacher. While prior studies aim to exp…
Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty
Yeseul Cho, Baekrok Shin, Changmin Kang +1
Recent advances in deep learning rely heavily on massive datasets, leading to substantial storage and training costs. Dataset pruning aims to alleviate this demand by discarding re…