23 citations · 56 across the 7 of their papers we have counts for
20 papers
What Language Model to Train if You Have One Million GPU Hours?
Teven Le Scao, Thomas Wang, Daniel Hesslow +16
The crystallization of modeling methods around the Transformer architecture has been a boon for practitioners. Simple, well-motivated architectural variations can transfer across t…
SciFact-Open: Towards open-domain scientific claim verification
David Wadden, Kyle Lo, Bailey Kuehl +4
While research on scientific claim verification has led to the development of powerful systems that appear to approach human performance, these approaches have yet to be tested in…
Continued Pretraining for Better Zero- and Few-Shot Promptability
Zhaofeng Wu, Robert L. Logan, Pete Walsh +4
Recently introduced language model prompting methods can achieve high accuracy in zero- and few-shot settings while requiring few to no learned task-specific parameters. Neverthele…
What Language Model Architecture and Pretraining Objective Work Best for Zero-Shot Generalization?
Thomas Wang, Adam Roberts, Daniel Hesslow +5
Large pretrained Transformer language models have been shown to exhibit zero-shot generalization, i.e. they can perform a wide variety of tasks that they were not explicitly traine…
Staged Training for Transformer Language Models
Sheng Shen, Pete Walsh, Kurt Keutzer +3
The current standard approach to scaling transformer language models trains each model size from a different random initialization. As an alternative, we consider a staged training…
FLEX: Unifying Evaluation for Few-Shot NLP
Jonathan Bragg, Arman Cohan, Kyle Lo +1
Few-shot NLP research is highly active, yet conducted in disjoint research threads with evaluation suites that lack challenging-yet-realistic testing setups and fail to employ care…