The Pile: An 800GB Dataset of Diverse Text for Language Modeling
arXiv:2101.00027
Abstract
Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models. With this in mind, we present \textit{the Pile}: an 825 GiB English text corpus targeted at training large-scale language models. The Pile is constructed from 22 diverse high-quality subsets -- both existing and newly constructed -- many of which derive from academic or professional sources. Our evaluation of the untuned performance of GPT-2 and GPT-3 on the Pile shows that these models struggle on many of its components, such as academic writing. Conversely, models trained on the Pile improve significantly over both Raw CC and CC-100 on all components of the Pile, while improving performance on downstream evaluations. Through an in-depth exploratory analysis, we document potentially concerning aspects of the data for prospective users. We make publicly available the code used in its construction.
References in corpus (11)
- Language Models are Few-Shot Learners
- NLTK: The Natural Language Toolkit
- Scaling Laws for Neural Language Models
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data
- Scaling Laws for Autoregressive Generative Modeling
- Analysing Mathematical Reasoning Abilities of Neural Models
- Generating Wikipedia by Summarizing Long Sequences
- Large image datasets: A pyrrhic win for computer vision?
- ABOUT ML: Annotation and Benchmarking on Understanding and Transparency of Machine Learning Lifecycles
- MT-Adapted Datasheets for Datasets: Template and Repository
Cited by in corpus (53)
- On the Opportunities and Risks of Foundation Models
- Evaluating Large Language Models Trained on Code
- Practical Program Repair in the Era of Large Pre-trained Language Models
- Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey
- Multimodal datasets: misogyny, pornography, and malignant stereotypes
- Natural Language Generation and Understanding of Big Code for AI-Assisted Programming: A Review
- Opening up ChatGPT: Tracking openness, transparency, and accountability in instruction-tuned text generators
- Materials science in the era of large language models: a perspective
- Better Together? An Evaluation of AI-Supported Code Translation
- The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
- ThoughtSource: A central hub for large language model reasoning data
- Few-Shot Bot: Prompt-Based Learning for Dialogue Systems
- Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
- Data Governance in the Age of Large-Scale Data-Driven Language Technology
- Quo Vadis ChatGPT? From Large Language Models to Large Knowledge Models
- "I'm fully who I am": Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation
- Several categories of Large Language Models (LLMs): A Short Survey
- An Evaluation on Large Language Model Outputs: Discourse and Memorization
- Large Language Models Can be Lazy Learners: Analyze Shortcuts in In-Context Learning
- A General Language Assistant as a Laboratory for Alignment
- Evaluating Biased Attitude Associations of Language Models in an Intersectional Context
- Evaluating Generative Patent Language Models
- Large Language Models, scientific knowledge and factuality: A framework to streamline human expert evaluation
- Language Modelling with Pixels
- The AI Community Building the Future? A Quantitative Analysis of Development Activity on Hugging Face Hub
- Intersectional Bias in Causal Language Models
- Tucano: Advancing Neural Text Generation for Portuguese
- ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models
- FABULA: Intelligence Report Generation Using Retrieval-Augmented Narrative Construction
- Intersectional Inquiry, on the Ground and in the Algorithm
- Fourier Transformer: Fast Long Range Modeling by Removing Sequence Redundancy with FFT Operator
- Truthful AI: Developing and governing AI that does not lie
- Future Lens: Anticipating Subsequent Tokens from a Single Hidden State
- Reviewing Clinical Knowledge in Medical Large Language Models: Training and Beyond
- Perhaps PTLMs Should Go to School -- A Task to Assess Open Book and Closed Book QA
- Language Modeling using LMUs: 10x Better Data Efficiency or Improved Scaling Compared to Transformers
- Structural Similarities Between Language Models and Neural Response Measurements
- Cracks in The Stack: Hidden Vulnerabilities and Licensing Risks in LLM Pre-Training Datasets
- Nano: Nested Human-in-the-Loop Reward Learning for Few-shot Language Model Control
- Is neural language acquisition similar to natural? A chronological probing study
- RAMP: Retrieval and Attribute-Marking Enhanced Prompting for Attribute-Controlled Translation
- An Empirical Exploration in Quality Filtering of Text Data
- DSAP: Analyzing Bias Through Demographic Comparison of Datasets
- Phrase-BERT: Improved Phrase Embeddings from BERT with an Application to Corpus Exploration
- Emergent inabilities? Inverse scaling over the course of pretraining
- Short paper: Models in the dark -- Rectification and erasure under GDPR in ML supply chains
- A Hierarchical Neural Framework for Classification and its Explanation in Large Unstructured Legal Documents
- Neural Program Generation Modulo Static Analysis
- Cut the CARP: Fishing for zero-shot story evaluation
- Towards A Measure Of General Machine Intelligence
- Remember what you did so you know what to do next