1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2025
BTS: Harmonizing Specialized Experts into a Generalist LLM
Qizhen Zhang, Prajjwal Bhargava, Chloe Bi +9
We present Branch-Train-Stitch (BTS), an efficient and flexible training algorithm for combining independently trained large language model (LLM) experts into a single, capable gen…
cs.CL2025★ 1 cited
Optimizing Pretraining Data Mixtures with LLM-Estimated Utility
William Held, Bhargavi Paranjape, Punit Singh Koura +3
Large Language Models improve with increasing amounts of high-quality training data. However, leveraging larger datasets requires balancing quality, quantity, and diversity across…