6 papers
The Effect of Training Task Diversity on In-Context Learning through the Lens of Low-Dimensional Subspaces
Soo Min Kwon, Alec S. Xu, Can Yaras +3
The transformer's emergent ability to perform in-context learning (ICL) has sparked a wide range of studies designed to understand its underlying mechanisms. Existing works often s…
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization
Soo Min Kwon, Ziteng Sun, Ananda Theertha Suresh +2
Group Relative Policy Optimization (GRPO) has emerged as a powerful algorithm for improving the reasoning capabilities of language models, but often fails to improve small models d…
Out-of-Distribution Generalization of In-Context Learning: A Low-Dimensional Subspace Perspective
Soo Min Kwon, Alec S. Xu, Can Yaras +2
The transformer's remarkable ability to perform in-context learning (ICL) has sparked a wide range of studies designed to understand its strengths and limitations. However, a theor…
An Overview of Low-Rank Structures in the Training and Adaptation of Large Models
Laura Balzano, Tianjiao Ding, Benjamin D. Haeffele +5
The substantial computational demands of modern large-scale deep learning present significant challenges for efficient training and deployment. Recent research has revealed a wides…
Decoupled Data Consistency with Diffusion Purification for Image Restoration
Xiang Li, Soo Min Kwon, Shijun Liang +3
Diffusion models have recently gained traction as a powerful class of deep generative priors, excelling in a wide range of image restoration tasks due to their exceptional ability…
Learning Dynamics of Deep Linear Networks Beyond the Edge of Stability
Avrajit Ghosh, Soo Min Kwon, Rongrong Wang +2
Deep neural networks trained using gradient descent with a fixed learning rate often operate in the regime of "edge of stability" (EOS), where the largest eigenvalue of the He…