activity
20202022
most citedThe Pile: An 800GB Dataset of Diverse Text for Language Modeling

492 citations · 603 across the 7 of their papers we have counts for

collaborators

8 papers

cs.LG202237 cited

Scaling Laws for Reward Model Overoptimization

Leo Gao, John Schulman, Jacob Hilton

In reinforcement learning from human feedback, it is common to optimize against a reward model trained to predict human preferences. Because the reward model is an imperfect proxy,…

cs.CL20223 cited

EleutherAI: Going Beyond "Open Science" to "Science in the Open"

Jason Phang, Herbie Bradley, Leo Gao +2

Over the past two years, EleutherAI has established itself as a radically novel initiative aimed at both promoting open-source research and conducting research in a transparent, op…

cs.CL202269 cited

GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Sid Black, Stella Biderman, Eric Hallahan +14

We introduce GPT-NeoX-20B, a 20 billion parameter autoregressive language model trained on the Pile, whose weights will be made freely and openly available to the public through a…

cs.CL2022

Datasheet for the Pile

Stella Biderman, Kieran Bicheno, Leo Gao

This datasheet describes the Pile, a 825 GiB dataset of human-authored text compiled by EleutherAI for use in large-scale language modeling. The Pile is comprised of 22 different t…

cs.CL2021

Cut the CARP: Fishing for zero-shot story evaluation

Shahbuland Matiana, JR Smith, Ryan Teehan +4

Recent advances in large-scale language models (Raffel et al., 2019; Brown et al., 2020) have brought significant qualitative and quantitative improvements in machine-driven text g…

cs.CL20212 cited

An Empirical Exploration in Quality Filtering of Text Data

Leo Gao

While conventional wisdom suggests that more aggressively filtering data from low-quality sources like Common Crawl always monotonically improves the quality of training data, we f…