activity
20192025
most citedThe Pile: An 800GB Dataset of Diverse Text for Language Modeling

492 citations · 585 across the 10 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL20226 cited

What Language Model to Train if You Have One Million GPU Hours?

Teven Le Scao, Thomas Wang, Daniel Hesslow +16

The crystallization of modeling methods around the Transformer architecture has been a boon for practitioners. Simple, well-motivated architectural variations can transfer across t…

cs.CL20223 cited

EleutherAI: Going Beyond "Open Science" to "Science in the Open"

Jason Phang, Herbie Bradley, Leo Gao +2

Over the past two years, EleutherAI has established itself as a radically novel initiative aimed at both promoting open-source research and conducting research in a transparent, op…

cs.CL202269 cited

GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Sid Black, Stella Biderman, Eric Hallahan +14

We introduce GPT-NeoX-20B, a 20 billion parameter autoregressive language model trained on the Pile, whose weights will be made freely and openly available to the public through a…

cs.CL20222 cited

Documenting Geographically and Contextually Diverse Data Sources: The BigScience Catalogue of Language Data and Resources

Angelina McMillan-Major, Zaid Alyafeai, Stella Biderman +15

In recent years, large-scale data collection efforts have prioritized the amount of data collected in order to improve the modeling capabilities of large language models. This prio…

cs.CL2022

Datasheet for the Pile

Stella Biderman, Kieran Bicheno, Leo Gao

This datasheet describes the Pile, a 825 GiB dataset of human-authored text compiled by EleutherAI for use in large-scale language modeling. The Pile is comprised of 22 different t…

cs.CL2021

Cut the CARP: Fishing for zero-shot story evaluation

Shahbuland Matiana, JR Smith, Ryan Teehan +4

Recent advances in large-scale language models (Raffel et al., 2019; Brown et al., 2020) have brought significant qualitative and quantitative improvements in machine-driven text g…