4 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.CL2022★ 1 cited
Continued Pretraining for Better Zero- and Few-Shot Promptability
Zhaofeng Wu, Robert L. Logan, Pete Walsh +4
Recently introduced language model prompting methods can achieve high accuracy in zero- and few-shot settings while requiring few to no learned task-specific parameters. Neverthele…
cs.CL2022★ 4 cited
Staged Training for Transformer Language Models
Sheng Shen, Pete Walsh, Kurt Keutzer +3
The current standard approach to scaling transformer language models trains each model size from a different random initialization. As an alternative, we consider a staged training…