Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework
Thomson Yen, Andrew Wei Tung Siah, Haozhe Chen +3
Careful curation of data sources can significantly improve the performance of LLM pre-training, but predominant approaches rely heavily on intuition or costly trial-and-error, maki…
cs.LG2025
Architectural and Inferential Inductive Biases For Exchangeable Sequence Modeling
Daksh Mittal, Ang Li, Tzu-Ching Yen +2
Autoregressive models have emerged as a powerful framework for modeling exchangeable sequences - i.i.d. observations when conditioned on some latent factor - enabling direct modeli…