26 citations · 33 across the 4 of their papers we have counts for
2 papers
cs.CL2023★ 7 cited
Scaling Expert Language Models with Unsupervised Domain Discovery
Suchin Gururangan, Margaret Li, Mike Lewis +4
Large language models are typically trained densely: all parameters are updated with respect to all inputs. This requires synchronization of billions of parameters across thousands…
cs.CL2022★ 26 cited
Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models
Margaret Li, Suchin Gururangan, Tim Dettmers +4
We present Branch-Train-Merge (BTM), a communication-efficient algorithm for embarrassingly parallel training of large language models (LLMs). We show it is possible to independent…