collaborators

8 papers

cs.LG2026

Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

Xinmu Ge, Zizhuo Zhang, Yu Huang +9

On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. It is commonly believed to enable the student model to distill knowledg…

cs.LG2026

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics

Yu Huang, Zixin Wen, Yuejie Chi +4

Reinforcement learning with verifiable rewards (RLVR) has been a main driver of recent breakthroughs in large reasoning models. Yet it remains a mystery how rewards based solely on…

cs.LG2026

Preconditioning Benefits of Spectral Orthogonalization in Muon

Jianhao Ma, Yu Huang, Yuejie Chi +1

The Muon optimizer, a matrix-structured algorithm that leverages spectral orthogonalization of gradients, is a milestone in the pretraining of large language models. However, the u…

cs.LG2025

Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization

Yu Huang, Zixin Wen, Aarti Singh +2

The ability to reason lies at the core of artificial intelligence (AI), and challenging problems usually call for deeper and longer reasoning to tackle. A crucial question about AI…

cs.LG2025

Dual-granularity Sinkhorn Distillation for Enhanced Learning from Long-tailed Noisy Data

Feng Hong, Yu Huang, Zihua Zhao +5

Real-world datasets for deep learning frequently suffer from the co-occurring challenges of class imbalance and label noise, hindering model performance. While methods exist for ea…

cs.LG2025

Transformers Meet In-Context Learning: A Universal Approximation Theory

Gen Li, Yuchen Jiao, Yu Huang +2

Large language models are capable of in-context learning, the ability to perform new tasks at test time using a handful of input-output examples, without parameter updates. We deve…