Publications (91)
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
Joel Niklaus, Atsuki Yamaguchi, Michal Štefánik +9
Knowledge is a Region in Weight Space for Fine-tuned Language Models
Almog Gueta, Elad Venezian, Colin Raffel +3
Emergent Abilities of Large Language Models
Jason Wei, Yi Tay, Rishi Bommasani +13
Feed-Forward Networks with Attention Can Solve Some Long-Term Memory Problems
Colin Raffel, Daniel P. W. Ellis
Petals: Collaborative Inference and Fine-tuning of Large Models
Alexander Borzunov, Dmitry Baranchuk, Tim Dettmers +5
Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model
Haikang Deng, Colin Raffel
Extracting Training Data from Large Language Models
Nicholas Carlini, Florian Tramer, Eric Wallace +9
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts +6
Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language
Zhenlin Xu, Marc Niethammer, Colin Raffel
Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning
Haokun Liu, Derek Tam, Mohammed Muqeeth +4
Deflecting Adversarial Attacks
Yao Qin, Nicholas Frosst, Colin Raffel +2
Realistic Evaluation of Deep Semi-Supervised Learning Algorithms
Avital Oliver, Augustus Odena, Colin Raffel +2
Multitask Prompted Training Enables Zero-Shot Task Generalization
Victor Sanh, Albert Webson, Colin Raffel +38
Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky +8
Do Transformer Modifications Transfer Across Implementations and Applications?
Sharan Narang, Hyung Won Chung, Yi Tay +13
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Loubna Ben Allal, Anton Lozhkov, Elie Bakouch +19
Crosslingual Generalization through Multitask Finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika +16
An Empirical Survey of Data Augmentation for Limited Data Learning in NLP
Jiaao Chen, Derek Tam, Colin Raffel +2
Theano: A Python framework for fast computation of mathematical expressions
The Theano Development Team, Rami Al-Rfou, Guillaume Alain +110
Efficiently Estimating Data Efficiency for Language Model Fine-tuning
Gyung Hyun Je, Colin Raffel
Git-Theta: A Git Extension for Collaborative Development of Machine Learning Models
Nikhil Kandpal, Brian Lester, Mohammed Muqeeth +6
Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
Jonathan Shen, Patrick Nguyen, Yonghui Wu +88
Uncovering Model Processing Strategies with Non-Negative Per-Example Fisher Factorization
Michael Matena, Colin Raffel
What Language Model Architecture and Pretraining Objective Work Best for Zero-Shot Generalization?
Thomas Wang, Adam Roberts, Daniel Hesslow +5
Online and Linear-Time Attention by Enforcing Monotonic Alignments
Colin Raffel, Minh-Thang Luong, Peter J. Liu +2
Large Language Models Struggle to Learn Long-Tail Knowledge
Nikhil Kandpal, Haikang Deng, Adam Roberts +2
ReMixMatch: Semi-Supervised Learning with Distribution Alignment and Augmentation Anchoring
David Berthelot, Nicholas Carlini, Ekin D. Cubuk +4
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
Bowen Pan, Yikang Shen, Haokun Liu +5
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Guilherme Penedo, Hynek KydlÃÄek, Loubna Ben allal +5
Realistic Evaluation of Model Merging for Compositional Generalization
Derek Tam, Yash Kant, Brian Lester +2
A Hierarchical Latent Vector Model for Learning Long-Term Structure in Music
Adam Roberts, Jesse Engel, Colin Raffel +2
Distributed Inference and Fine-tuning of Large Language Models Over The Internet
Alexander Borzunov, Max Ryabinin, Artem Chumachenko +5
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts +5
On Training Sample Memorization: Lessons from Benchmarking Generative Modeling with a Large-scale Competition
Ching-Yuan Bai, Hsuan-Tien Lin, Colin Raffel +1
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
YuXin Li, Felix Dangel, Derek Tam +1
WT5?! Training Text-to-Text Models to Explain their Predictions
Sharan Narang, Colin Raffel, Katherine Lee +3
Poker-CNN: A Pattern Learning Strategy for Making Draws and Bets in Poker Games
Nikolai Yakovenko, Liangliang Cao, Colin Raffel +1
Robust and Generalizable Visual Representation Learning via Random Convolutions
Zhenlin Xu, Deyi Liu, Junlin Yang +2
Monotonic Chunkwise Attention
Chung-Cheng Chiu, Colin Raffel
Onsets and Frames: Dual-Objective Piano Transcription
Curtis Hawthorne, Erich Elsen, Jialin Song +6
Deduplicating Training Data Mitigates Privacy Risks in Language Models
Nikhil Kandpal, Eric Wallace, Colin Raffel
MixMatch: A Holistic Approach to Semi-Supervised Learning
David Berthelot, Nicholas Carlini, Ian Goodfellow +3
Top-k Training of GANs: Improving GAN Performance by Throwing Away Bad Samples
Samarth Sinha, Zhengli Zhao, Anirudh Goyal +2
Understanding and Improving Interpolation in Autoencoders via an Adversarial Regularizer
David Berthelot, Colin Raffel, Aurko Roy +1
Efficient Methods for Natural Language Processing: A Survey
Marcos Treviso, Ji-Ung Lee, Tianchu Ji +19
ColD Fusion: Collaborative Descent for Distributed Multitask Finetuning
Shachar Don-Yehiya, Elad Venezian, Colin Raffel +3
Training a Subsampling Mechanism in Expectation
Colin Raffel, Dieterich Lawson
PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts
Stephen H. Bach, Victor Sanh, Zheng-Xin Yong +24
How Much Knowledge Can You Pack Into the Parameters of a Language Model?
Adam Roberts, Colin Raffel, Noam Shazeer
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
Prateek Yadav, Leshem Choshen, Colin Raffel +1
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution
Fengyuan Liu, Nikhil Kandpal, Colin Raffel
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
BigScience Workshop, :, Teven Le Scao +391
Is Generator Conditioning Causally Related to GAN Performance?
Augustus Odena, Jacob Buckman, Catherine Olsson +4
ByT5: Towards a token-free future with pre-trained byte-to-byte models
Linting Xue, Aditya Barua, Noah Constant +5
Scaling Data-Constrained Language Models
Niklas Muennighoff, Alexander M. Rush, Boaz Barak +6
Monotonic Infinite Lookback Attention for Simultaneous Machine Translation
Naveen Arivazhagan, Colin Cherry, Wolfgang Macherey +5
Imperceptible, Robust, and Targeted Adversarial Examples for Automatic Speech Recognition
Yao Qin, Nicholas Carlini, Ian Goodfellow +2
Improving Few-Shot Generalization by Exploring and Exploiting Auxiliary Data
Alon Albalak, Colin Raffel, William Yang Wang
Position: The Most Expensive Part of an LLM should be its Training Data
Nikhil Kandpal, Colin Raffel
Learning to Route Among Specialized Experts for Zero-Shot Generalization
Mohammed Muqeeth, Haokun Liu, Yufan Liu +1
Learning a Latent Space of Multitrack Measures
Ian Simon, Adam Roberts, Colin Raffel +3
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
Prateek Yadav, Colin Raffel, Mohammed Muqeeth +6
What Language Model to Train if You Have One Million GPU Hours?
Teven Le Scao, Thomas Wang, Daniel Hesslow +16
Towards GAN Benchmarks Which Require Generalization
Ishaan Gulrajani, Colin Raffel, Luke Metz
The Appeal and Reality of Recycling LoRAs with Adaptive Merging
Haokun Liu, Gyung Hyun Je, Marco Ciccone +3
TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior
Gül Sena AltıntaÅ, Malikeh Ehghaghi, Brian Lester +4
Model Merging via Data-Free Covariance Estimation
Marawan Gamal Abdel Hameed, Derek Tam, Pascal Jr Tikeng Notsawo +2
Merging Models with Fisher-Weighted Averaging
Michael Matena, Colin Raffel
Efficient Online Data Mixing For Language Model Pre-Training
Alon Albalak, Liangming Pan, Colin Raffel +1
Soft Merging of Experts with Adaptive Routing
Mohammed Muqeeth, Haokun Liu, Colin Raffel
Bidirectional Language Models Are Also Few-shot Learners
Ajay Patel, Bryan Li, Mohammad Sadegh Rasooli +3
Training Neural Networks with Fixed Sparse Masks
Yi-Lin Sung, Varun Nair, Colin Raffel
Improving and Simplifying Pattern Exploiting Training
Derek Tam, Rakesh R Menon, Mohit Bansal +2
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
Ajay Patel, Colin Raffel, Chris Callison-Burch
Evaluating the Factual Consistency of Large Language Models Through News Summarization
Derek Tam, Anisha Mascarenhas, Shiyue Zhang +3
The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions
Devin Kwok, Gül Sena AltıntaÅ, Colin Raffel +1
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Guilherme Penedo, Hynek KydlÃÄek, Vinko SabolÄec +7
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Nikhil Kandpal, Brian Lester, Colin Raffel +24
Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models
Malikeh Ehghaghi, Boglárka Ecsedi, Marsha Chechik +1
Enhancing Training Data Attribution with Representational Optimization
Weiwei Sun, Haokun Liu, Nikhil Kandpal +2
NeurIPS 2020 EfficientQA Competition: Systems, Analyses and Lessons Learned
Sewon Min, Jordan Boyd-Graber, Chris Alberti +50
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
Ajay Patel, Colin Raffel, Chris Callison-Burch
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao +448
Learning Hard Alignments with Variational Inference
Dieterich Lawson, Chung-Cheng Chiu, George Tucker +3
A Combinatorial Perspective on the Optimization of Shallow ReLU Networks
Michael Matena, Colin Raffel
Merging by Matching Models in Task Parameter Subspaces
Derek Tam, Mohit Bansal, Colin Raffel
Detecting and Diagnosing Adversarial Images with Class-Conditional Capsule Reconstructions
Yao Qin, Nicholas Frosst, Sara Sabour +3
TIES-Merging: Resolving Interference When Merging Models
Prateek Yadav, Derek Tam, Leshem Choshen +2
Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$
Adam Roberts, Hyung Won Chung, Anselm Levskaya +40
A Survey on Data Selection for Language Models
Alon Albalak, Yanai Elazar, Sang Michael Xie +11
FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence
Kihyuk Sohn, David Berthelot, Chun-Liang Li +6