From the 1 of 7 linked papers with an AI index.
7 papers
Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting
Xiang Yuan, Kaiqing Lei, Zhenyu Jin +3
The paper proposes a Bayesian method that learns optimal domain weights for multi‑domain pre‑training of large language models by inferring a Dirichlet distribution with Gamma prio…
Sharper Analysis of Single-Loop Methods for Bilevel Optimization
Yubo Zhou, Jun Shu, Luo Luo +4
Bilevel optimization underpins many machine learning applications, including hyperparameter optimization, meta-learning, neural architecture search, and reinforcement learning. Whi…
A Limit Theory of Foundation Models: A Mathematical Approach to Understanding Emergent Intelligence and Scaling Laws
Jun Shu, Junxiong Jia, Deyu Meng +1
Emergent intelligence have played a major role in the modern AI development. While existing studies primarily rely on empirical observations to characterize this phenomenon, a rigo…
Understanding the Generalization of Bilevel Programming in Hyperparameter Optimization: A Tale of Bias-Variance Decomposition
Yubo Zhou, Jun Shu, Junmin Liu +1
Gradient-based hyperparameter optimization (HPO) have emerged recently, leveraging bilevel programming techniques to optimize hyperparameter by estimating hypergradient w.r.t. vali…
KoopGen: Koopman Generator Networks for Representing and Predicting Dynamical Systems with Continuous Spectra
Liangyu Su, Jun Shu, Rui Liu +2
Representing and predicting high-dimensional and spatiotemporally chaotic dynamical systems remains a fundamental challenge in dynamical systems and machine learning. Although data…
Feed Two Birds with One Scone: Exploiting Function-Space Regularization for Both OOD Robustness and ID Fine-Tuning Performance
Xiang Yuan, Jun Shu, Deyu meng +1
Robust fine-tuning aims to achieve competitive in-distribution (ID) performance while maintaining the out-of-distribution (OOD) robustness of a pre-trained model when transferring…