3k citations · 5k across the 3 of their papers we have counts for
3 papers
cs.CL2020★ 3k cited
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder +28
Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typicall…
cs.LG2020★ 1.5k cited
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan +7
We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute us…
cs.LG2019★ 484 cited
Generating Long Sequences with Sparse Transformers
Rewon Child, Scott Gray, Alec Radford +1
Transformers are powerful sequence models, but require time and memory that grows quadratically with the sequence length. In this paper we introduce sparse factorizations of the at…