activity
20182025
most citedBeyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

565 citations · 1.5k across the 23 of their papers we have counts for

collaborators
Showing 2022Show all

5 papers · 1 filter

cs.CL2022★ 6 cited

Transcending Scaling Laws with 0.1% Extra Compute

Yi Tay, Jason Wei, Hyung Won Chung +13

Scaling language models improves performance but comes with significant computational costs. This paper proposes UL2R, a method that substantially improves existing language models…

cs.CL2022★ 3 cited

Knowledge Prompts: Injecting World Knowledge into Language Models through Soft Prompts

Cicero Nogueira dos Santos, Zhe Dong, Daniel Cer +4

Soft prompts have been recently proposed as a tool for adapting large frozen language models (LMs) to new tasks. In this work, we repurpose soft prompts to the task of injecting wo…

cs.CL2022★ 6 cited

Reducing Retraining by Recycling Parameter-Efficient Prompts

Brian Lester, Joshua Yurtsever, Siamak Shakeri +1

Parameter-efficient methods are able to use a single frozen pre-trained large language model (LLM) to perform many tasks by learning task-specific soft prompts that modulate model…

cs.CL2022★ 565 cited

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao +448

Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabil…

cs.CL2022★ 6 cited

Counterfactual Data Augmentation improves Factuality of Abstractive Summarization

Dheeraj Rajagopal, Siamak Shakeri, Cicero Nogueira dos Santos +2

Abstractive summarization systems based on pretrained language models often generate coherent but factually inconsistent sentences. In this paper, we present a counterfactual data…