Publications (72)
Pitfalls in Machine Learning Research: Reexamining the Development Cycle
Stella Biderman, Walter J. Scheirer
Fooling MOSS Detection with Pretrained Language Models
Stella Biderman, Edward Raff
The Ghost in the Keys: A Disklavier Demo for Human-AI Musical Co-Creativity
Louis Bradshaw, Alexander Spangher, Stella Biderman +1
Holographic Global Convolutional Networks for Long-Range Prediction Tasks in Malware Detection
Mohammad Mahmudul Alam, Edward Raff, Stella Biderman +2
Data Governance in the Age of Large-Scale Data-Driven Language Technology
Yacine Jernite, Huu Nguyen, Stella Biderman +18
BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing
Jason Alan Fries, Leon Weber, Natasha Seelam +40
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Mubashara Akhtar, Anka Reuel, Prajna Soni +34
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
Avijit Ghosh, Anka Reuel, Jenny Chim +45
Multitask Prompted Training Enables Zero-Shot Task Generalization
Victor Sanh, Albert Webson, Colin Raffel +38
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
Julia Kreutzer, Isaac Caswell, Lisa Wang +49
Cut the CARP: Fishing for zero-shot story evaluation
Shahbuland Matiana, JR Smith, Ryan Teehan +4
Crosslingual Generalization through Multitask Finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika +16
PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs
Oskar van der Wal, Pietro Lesci, Max Muller-Eberstein +4
Emergent and Predictable Memorization in Large Language Models
Stella Biderman, USVSN Sai Prashanth, Lintang Sutawika +4
Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence
Bo Peng, Daniel Goldstein, Quentin Anthony +27
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45
Recasting Self-Attention with Holographic Reduced Representations
Mohammad Mahmudul Alam, Edward Raff, Stella Biderman +2
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
Stay on topic with Classifier-Free Guidance
Guillaume Sanchez, Honglu Fan, Alexander Spangher +3
Can Transformers Learn to Solve Problems Recursively?
Shizhuo Dylan Zhang, Curt Tigges, Stella Biderman +2
Towards Best Practices for Open Datasets for LLM Training
Stefan Baack, Stella Biderman, Kasia Odrozek +36
Grokking Group Multiplication with Cosets
Dashiell Stander, Qinan Yu, Honglu Fan +1
Towards a Formal Model of Narratives
Louis Castricato, Stella Biderman, Rogelio E. Cardona-Rivera +1
A Walsh Hadamard Derived Linear Vector Symbolic Architecture
Mohammad Mahmudul Alam, Alexander Oberle, Edward Raff +3
On the Sensitivity of k-Uniform Hypergraph Properties
Stella Biderman, Kevin Cuddy, Ang Li +1
Bridging the Data Provenance Gap Across Text, Speech and Video
Shayne Longpre, Nikhil Singh, Manuel Cherep +40
Magic: The Gathering is Turing Complete
Alex Churchill, Stella Biderman, Austin Herrick
The Responsible Foundation Model Development Cheatsheet: A Review of Tools & Resources
Shayne Longpre, Stella Biderman, Alon Albalak +20
Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
Stella Biderman, Mohammad Aflah Khan, Niloofar Mireshghallah +3
BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting
Zheng-Xin Yong, Hailey Schoelkopf, Niklas Muennighoff +12
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang +51
Bergson: An Open Source Library for Data Attribution
Lucia Quirke, Louis Jaburi, David Johnston +6
Eliciting Latent Predictions from Transformers with the Tuned Lens
Nora Belrose, Igor Ostrovsky, Lev McKinney +5
Scaling Self-Supervised Representation Learning for Symbolic Piano Performance
Louis Bradshaw, Honglu Fan, Alexander Spangher +2
On the Societal Impact of Open Foundation Models
Sayash Kapoor, Rishi Bommasani, Kevin Klyman +22
RWKV: Reinventing RNNs for the Transformer Era
Bo Peng, Eric Alcaide, Quentin Anthony +31
The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs
Laura Ruis, Akbir Khan, Stella Biderman +3
LLM Circuit Analyses Are Consistent Across Training and Scale
Curt Tigges, Michael Hanna, Qinan Yu +1
Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon
USVSN Sai Prashanth, Alvin Deng, Kyle O'Brien +9
EleutherAI: Going Beyond "Open Science" to "Science in the Open"
Jason Phang, Herbie Bradley, Leo Gao +2
Lessons from the Trenches on Reproducible Evaluation of Language Models
Stella Biderman, Hailey Schoelkopf, Lintang Sutawika +27
Suppressing Pink Elephants with Direct Principle Feedback
Louis Castricato, Nathan Lile, Suraj Anand +3
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
BigScience Workshop, :, Teven Le Scao +391
VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance
Katherine Crowson, Stella Biderman, Daniel Kornis +4
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
Rylan Schaeffer, Hailey Schoelkopf, Brando Miranda +6
Adversarial Samples Are Not Created Equal
Jennifer Crawford, Amol Khanna, Fred Lu +4
LEACE: Perfect linear concept erasure in closed form
Nora Belrose, David Schneider-Joseph, Shauli Ravfogel +3
Explaining and Mitigating Crosslingual Tokenizer Inequities
Catherine Arnett, Tyler A. Chang, Stella Biderman +1
What Language Model to Train if You Have One Million GPU Hours?
Teven Le Scao, Thomas Wang, Daniel Hesslow +16
Open Problems in Mechanistic Interpretability
Lee Sharkey, Bilal Chughtai, Joshua Batson +26
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
Guijin Son, Jiwoo Hong, Honglu Fan +8
GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Sid Black, Stella Biderman, Eric Hallahan +14
Documenting Geographically and Contextually Diverse Data Sources: The BigScience Catalogue of Language Data and Resources
Angelina McMillan-Major, Zaid Alyafeai, Stella Biderman +15
Consent in Crisis: The Rapid Decline of the AI Data Commons
Shayne Longpre, Robert Mahari, Ariel Lee +46
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
Kyle O'Brien, Stephen Casper, Quentin Anthony +7
Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns
Max Kamachee, Stephen Casper, Michelle L. Ding +4
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Nikhil Kandpal, Brian Lester, Colin Raffel +24
GAIA Search: Hugging Face and Pyserini Interoperability for NLP Training Data Exploration
Aleksandra Piktus, Odunayo Ogundepo, Christopher Akiki +6
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao +448
Magic: the Gathering is as Hard as Arithmetic
Stella Biderman
Datasheet for the Pile
Stella Biderman, Kieran Bicheno, Leo Gao
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Leo Gao, Stella Biderman, Sid Black +9
LoRA-Muon: Spectral Steepest Descent on the Low-Rank Manifold
Franz Louis Cesista, Katherine Crowson, Cédric Simal +1
Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving
Pawan Sasanka Ammanamanchi, Siddharth Bhat, Stella Biderman
Quantifying the Effect of Test Set Contamination on Generative Evaluations
Rylan Schaeffer, Joshua Kazdan, Baber Abbasi +8
The Case for Co-Designing Model Architectures with Hardware
Quentin Anthony, Jacob Hatef, Deepak Narayanan +6
Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
Anka Reuel, Avijit Ghosh, Jenny Chim +32
Llemma: An Open Language Model For Mathematics
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster +6
KMMLU: Measuring Massive Multitask Language Understanding in Korean
Guijin Son, Hanwool Lee, Sungdong Kim +6
Transformer-Based Models Are Not Yet Perfect At Learning to Emulate Structural Recursion
Dylan Zhang, Curt Tigges, Zory Zhang +3
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Stella Biderman, Hailey Schoelkopf, Quentin Anthony +10
Capability Provenance in Language Models: A Case Study in Social Reasoning
Glenn Matlin, Chandreyi Chakraborty, Saehee Eom +8