3 papers
cs.LG2026
Mixtera: A Data Plane for Foundation Model Training
Maximilian Böther, Xiaozhe Yao, Tolga Kerimoglu +3
State-of-the-art large language and vision models are trained over trillions of tokens that are aggregated from a large variety of sources. As training data collections grow, manua…
cs.LG2025
On Distributed Larger-Than-Memory Subset Selection With Pairwise Submodular Functions
Maximilian Böther, Abraham Sebastian, Pranjal Awasthi +2
Modern datasets span billions of samples, making training on all available data infeasible. Selecting a high quality subset helps in reducing training costs and enhancing model qua…
cs.LG2025
Modyn: Data-Centric Machine Learning Pipeline Orchestration
Maximilian Böther, Ties Robroek, Viktor Gsteiger +4
In real-world machine learning (ML) pipelines, datasets are continuously growing. Models must incorporate this new training data to improve generalization and adapt to potential di…