129 citations · 164 across the 4 of their papers we have counts for
5 papers · 1 filter
Long-form factuality in large language models
Jerry Wei, Chengrun Yang, Xinying Song +9
Large language models (LLMs) often generate content that contains factual errors when responding to fact-seeking prompts on open-ended topics. To benchmark a model's long-form fact…
Simple synthetic data reduces sycophancy in large language models
Jerry Wei, Da Huang, Yifeng Lu +2
Sycophancy is an undesirable behavior where models tailor their responses to follow a human user's view even when that view is not objectively correct (e.g., adapting liberal views…
DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining
Sang Michael Xie, Hieu Pham, Xuanyi Dong +7
The mixture proportions of pretraining data domains (e.g., Wikipedia, books, web text) greatly affect language model (LM) performance. In this paper, we propose Domain Reweighting…
Symbol tuning improves in-context learning in language models
Jerry Wei, Le Hou, Andrew Lampinen +8
We present symbol tuning - finetuning language models on in-context input-label pairs where natural language labels (e.g., "positive/negative sentiment") are replaced with arbitrar…
SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network
William Chan, Daniel Park, Chris Lee +3
We present SpeechStew, a speech recognition model that is trained on a combination of various publicly available speech recognition datasets: AMI, Broadcast News, Common Voice, Lib…