2 papers
cs.LG2025
On Distributed Larger-Than-Memory Subset Selection With Pairwise Submodular Functions
Maximilian Böther, Abraham Sebastian, Pranjal Awasthi +2
Modern datasets span billions of samples, making training on all available data infeasible. Selecting a high quality subset helps in reducing training costs and enhancing model qua…
cs.LG2024
Majority Kernels: An Approach to Leverage Big Model Dynamics for Efficient Small Model Training
Hanna Mazzawi, Pranjal Awasthi, Xavi Gonzalvo +1
Recent breakthroughs and successful deployment of large language and vision models in a constrained environment predominantly follow a two phase approach. First, large models are t…