bayesian methods 1data mixture optimization 1domain weighting 1large language models 1pretraining data selection 1
From the 1 of 4 linked papers with an AI index.
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting
Xiang Yuan, Kaiqing Lei, Zhenyu Jin +3
The paper proposes a Bayesian method that learns optimal domain weights for multi‑domain pre‑training of large language models by inferring a Dirichlet distribution with Gamma prio…
cs.LG2025
Feed Two Birds with One Scone: Exploiting Function-Space Regularization for Both OOD Robustness and ID Fine-Tuning Performance
Xiang Yuan, Jun Shu, Deyu meng +1
Robust fine-tuning aims to achieve competitive in-distribution (ID) performance while maintaining the out-of-distribution (OOD) robustness of a pre-trained model when transferring…
cs.LG2025
Improving Memory Efficiency for Training KANs via Meta Learning
Zhangchi Zhao, Jun Shu, Deyu Meng +1
Inspired by the Kolmogorov-Arnold representation theorem, KANs offer a novel framework for function approximation by replacing traditional neural network weights with learnable uni…