26 citations · 36 across the 4 of their papers we have counts for
4 papers
Predicting vs. Acting: A Trade-off Between World Modeling & Agent Modeling
Margaret Li, Weijia Shi, Artidoro Pagnoni +2
RLHF-aligned LMs have shown unprecedented ability on both benchmarks and long-form text generation, yet they struggle with one foundational task: next-token prediction. As RLHF mod…
A Taxonomy of Ambiguity Types for NLP
Margaret Y. Li, Alisa Liu, Zhaofeng Wu +1
Ambiguity is an critical component of language that allows for more effective communication between speakers, but is often ignored in NLP. Recent work suggests that NLP systems may…
Scaling Expert Language Models with Unsupervised Domain Discovery
Suchin Gururangan, Margaret Li, Mike Lewis +4
Large language models are typically trained densely: all parameters are updated with respect to all inputs. This requires synchronization of billions of parameters across thousands…
Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models
Margaret Li, Suchin Gururangan, Tim Dettmers +4
We present Branch-Train-Merge (BTM), a communication-efficient algorithm for embarrassingly parallel training of large language models (LLMs). We show it is possible to independent…