collaborators

7 papers

cs.LG2026

Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting

Xiang Yuan, Kaiqing Lei, Zhenyu Jin +3

The paper proposes a Bayesian method that learns optimal domain weights for multi‑domain pre‑training of large language models by inferring a Dirichlet distribution with Gamma prio…

cs.LG2026

Sharper Analysis of Single-Loop Methods for Bilevel Optimization

Yubo Zhou, Jun Shu, Luo Luo +4

Bilevel optimization underpins many machine learning applications, including hyperparameter optimization, meta-learning, neural architecture search, and reinforcement learning. Whi…

cs.LG2026

A Limit Theory of Foundation Models: A Mathematical Approach to Understanding Emergent Intelligence and Scaling Laws

Jun Shu, Junxiong Jia, Deyu Meng +1

Emergent intelligence have played a major role in the modern AI development. While existing studies primarily rely on empirical observations to characterize this phenomenon, a rigo…

cs.LG2026

Understanding the Generalization of Bilevel Programming in Hyperparameter Optimization: A Tale of Bias-Variance Decomposition

Yubo Zhou, Jun Shu, Junmin Liu +1

Gradient-based hyperparameter optimization (HPO) have emerged recently, leveraging bilevel programming techniques to optimize hyperparameter by estimating hypergradient w.r.t. vali…

cs.LG2026

KoopGen: Koopman Generator Networks for Representing and Predicting Dynamical Systems with Continuous Spectra

Liangyu Su, Jun Shu, Rui Liu +2

Representing and predicting high-dimensional and spatiotemporally chaotic dynamical systems remains a fundamental challenge in dynamical systems and machine learning. Although data…

cs.LG2025

Feed Two Birds with One Scone: Exploiting Function-Space Regularization for Both OOD Robustness and ID Fine-Tuning Performance

Xiang Yuan, Jun Shu, Deyu meng +1

Robust fine-tuning aims to achieve competitive in-distribution (ID) performance while maintaining the out-of-distribution (OOD) robustness of a pre-trained model when transferring…