activity
20182022
most citedRealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

18 citations · 58 across the 8 of their papers we have counts for

collaborators

13 papers

cs.CL2025

BTS: Harmonizing Specialized Experts into a Generalist LLM

Qizhen Zhang, Prajjwal Bhargava, Chloe Bi +9

We present Branch-Train-Stitch (BTS), an efficient and flexible training algorithm for combining independently trained large language model (LLM) experts into a single, capable gen…

cs.LG20227 cited

lo-fi: distributed fine-tuning without communication

Mitchell Wortsman, Suchin Gururangan, Shen Li +4

When fine-tuning large neural networks, it is common to use multiple nodes and to communicate gradients at each optimization step. By contrast, we investigate completely local fine…

cs.CL2022

M2D2: A Massively Multi-domain Language Modeling Dataset

Machel Reid, Victor Zhong, Suchin Gururangan +1

We present M2D2, a fine-grained, massively multi-domain corpus for studying domain adaptation in language models (LMs). M2D2 consists of 8.5B tokens and spans 145 domains extracted…

cs.CL20226 cited

Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection

Suchin Gururangan, Dallas Card, Sarah K. Dreier +5

Language models increasingly rely on massive web dumps for diverse text data. However, these sources are rife with undesirable content. As such, resources like Wikipedia, books, an…

cs.CL2021

Expected Validation Performance and Estimation of a Random Variable's Maximum

Jesse Dodge, Suchin Gururangan, Dallas Card +2

Research in NLP is often supported by experimental results, and improved reporting of such results can lead to better understanding and more reproducible science. In this paper we…

cs.CL2021

DEMix Layers: Disentangling Domains for Modular Language Modeling

Suchin Gururangan, Mike Lewis, Ari Holtzman +2

We introduce a new domain expert mixture (DEMix) layer that enables conditioning a language model (LM) on the domain of the input text. A DEMix layer is a collection of expert feed…