activity
20212025
most citedCarbon Emissions and Large Neural Network Training

129 citations · 164 across the 4 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2024

Long-form factuality in large language models

Jerry Wei, Chengrun Yang, Xinying Song +9

Large language models (LLMs) often generate content that contains factual errors when responding to fact-seeking prompts on open-ended topics. To benchmark a model's long-form fact…

cs.CL2023

Simple synthetic data reduces sycophancy in large language models

Jerry Wei, Da Huang, Yifeng Lu +2

Sycophancy is an undesirable behavior where models tailor their responses to follow a human user's view even when that view is not objectively correct (e.g., adapting liberal views…

cs.CL2023

DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Sang Michael Xie, Hieu Pham, Xuanyi Dong +7

The mixture proportions of pretraining data domains (e.g., Wikipedia, books, web text) greatly affect language model (LM) performance. In this paper, we propose Domain Reweighting…

cs.CL2023

Symbol tuning improves in-context learning in language models

Jerry Wei, Le Hou, Andrew Lampinen +8

We present symbol tuning - finetuning language models on in-context input-label pairs where natural language labels (e.g., "positive/negative sentiment") are replaced with arbitrar…

cs.CL2021

SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network

William Chan, Daniel Park, Chris Lee +3

We present SpeechStew, a speech recognition model that is trained on a combination of various publicly available speech recognition datasets: AMI, Broadcast News, Common Voice, Lib…