collaborators

5 papers

cs.LG2026

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon

Huaqing Zhang, Jingchu Gai, Juno Kim +2

Online imitation learning (IL), particularly on-policy distillation, has emerged as a strong LLM post-training approach, often outperforming offline supervised fine-tuning (SFT). Y…

cs.LG2026

Beyond RLHF: A Unified Theoretical Framework of Alignment

Jihun Yun, Juno Kim, Jongho Park +4

Alignment via reinforcement learning from human feedback (RLHF) has become the dominant paradigm for controlling the quality of outputs from large language models (LLMs). However,…

stat.ML2026

Sharp Capacity Thresholds in Linear Associative Memory: From Top-1 Retrieval to Tail-Average Learning

Nicholas Barnfield, Juno Kim, Eshaan Nichani +2

How many key-value associations can a linear memory store? The answer depends not only on the degrees of freedom in the memory matrix, but also on the retrieval c…

cs.LG2026

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory

Juno Kim, Eshaan Nichani, Denny Wu +2

Spectral optimizers such as Muon have recently shown strong empirical performance in large-scale language model training, but the source and extent of their advantage remain poorly…

cs.LG2026

Coverage Improvement and Fast Convergence of On-policy Preference Learning

Juno Kim, Jihun Yun, Jason D. Lee +1

Online on-policy preference learning algorithms for language model alignment such as online direct policy optimization (DPO) can significantly outperform their offline counterparts…