NewEvery arXiv paper, its researchers & institutions — mapped.
papers

Publications (91)

cs.CL2026

How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data

Joel Niklaus, Atsuki Yamaguchi, Michal Štefánik +9

cs.LG2023

Knowledge is a Region in Weight Space for Fine-tuned Language Models

Almog Gueta, Elad Venezian, Colin Raffel +3

cs.CL2022

Emergent Abilities of Large Language Models

Jason Wei, Yi Tay, Rishi Bommasani +13

cs.LG2016

Feed-Forward Networks with Attention Can Solve Some Long-Term Memory Problems

Colin Raffel, Daniel P. W. Ellis

cs.LG2023

Petals: Collaborative Inference and Fine-tuning of Large Models

Alexander Borzunov, Dmitry Baranchuk, Tim Dettmers +5

cs.CL2024

Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model

Haikang Deng, Colin Raffel

cs.CR2021

Extracting Training Data from Large Language Models

Nicholas Carlini, Florian Tramer, Eric Wallace +9

cs.LG2023

Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Colin Raffel, Noam Shazeer, Adam Roberts +6

cs.LG2022

Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language

Zhenlin Xu, Marc Niethammer, Colin Raffel

cs.LG2022

Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning

Haokun Liu, Derek Tam, Mohammed Muqeeth +4

cs.LG2020

Deflecting Adversarial Attacks

Yao Qin, Nicholas Frosst, Colin Raffel +2

cs.LG2019

Realistic Evaluation of Deep Semi-Supervised Learning Algorithms

Avital Oliver, Augustus Odena, Colin Raffel +2

cs.LG2022

Multitask Prompted Training Enables Zero-Shot Task Generalization

Victor Sanh, Albert Webson, Colin Raffel +38

cs.CL2021

Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky +8

cs.LG2021

Do Transformer Modifications Transfer Across Implementations and Applications?

Sharan Narang, Hyung Won Chung, Yi Tay +13

cs.CL2025

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Loubna Ben Allal, Anton Lozhkov, Elie Bakouch +19

cs.CL2023

Crosslingual Generalization through Multitask Finetuning

Niklas Muennighoff, Thomas Wang, Lintang Sutawika +16

cs.CL2021

An Empirical Survey of Data Augmentation for Limited Data Learning in NLP

Jiaao Chen, Derek Tam, Colin Raffel +2

cs.SC2016

Theano: A Python framework for fast computation of mathematical expressions

The Theano Development Team, Rami Al-Rfou, Guillaume Alain +110

cs.LG2025

Efficiently Estimating Data Efficiency for Language Model Fine-tuning

Gyung Hyun Je, Colin Raffel

cs.LG2023

Git-Theta: A Git Extension for Collaborative Development of Machine Learning Models

Nikhil Kandpal, Brian Lester, Mohammed Muqeeth +6

cs.LG2019

Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

Jonathan Shen, Patrick Nguyen, Yonghui Wu +88

cs.LG2026

Uncovering Model Processing Strategies with Non-Negative Per-Example Fisher Factorization

Michael Matena, Colin Raffel

cs.CL2022

What Language Model Architecture and Pretraining Objective Work Best for Zero-Shot Generalization?

Thomas Wang, Adam Roberts, Daniel Hesslow +5

cs.LG2017

Online and Linear-Time Attention by Enforcing Monotonic Alignments

Colin Raffel, Minh-Thang Luong, Peter J. Liu +2

cs.CL2023

Large Language Models Struggle to Learn Long-Tail Knowledge

Nikhil Kandpal, Haikang Deng, Adam Roberts +2

cs.LG2020

ReMixMatch: Semi-Supervised Learning with Distribution Alignment and Augmentation Anchoring

David Berthelot, Nicholas Carlini, Ekin D. Cubuk +4

cs.LG2024

Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Bowen Pan, Yikang Shen, Haokun Liu +5

cs.CL2024

The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal +5

cs.LG2024

Realistic Evaluation of Model Merging for Compositional Generalization

Derek Tam, Yash Kant, Brian Lester +2

cs.LG2019

A Hierarchical Latent Vector Model for Learning Long-Term Structure in Music

Adam Roberts, Jesse Engel, Colin Raffel +2

cs.LG2023

Distributed Inference and Fine-tuning of Large Language Models Over The Internet

Alexander Borzunov, Max Ryabinin, Artem Chumachenko +5

cs.CL2021

mT5: A massively multilingual pre-trained text-to-text transformer

Linting Xue, Noah Constant, Adam Roberts +5

cs.LG2021

On Training Sample Memorization: Lessons from Benchmarking Generative Modeling with a Large-scale Competition

Ching-Yuan Bai, Hsuan-Tien Lin, Colin Raffel +1

cs.LG2025

Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator

YuXin Li, Felix Dangel, Derek Tam +1

cs.CL2020

WT5?! Training Text-to-Text Models to Explain their Predictions

Sharan Narang, Colin Raffel, Katherine Lee +3

cs.AI2015

Poker-CNN: A Pattern Learning Strategy for Making Draws and Bets in Poker Games

Nikolai Yakovenko, Liangliang Cao, Colin Raffel +1

cs.CV2021

Robust and Generalizable Visual Representation Learning via Random Convolutions

Zhenlin Xu, Deyi Liu, Junlin Yang +2

cs.CL2018

Monotonic Chunkwise Attention

Chung-Cheng Chiu, Colin Raffel

cs.SD2018

Onsets and Frames: Dual-Objective Piano Transcription

Curtis Hawthorne, Erich Elsen, Jialin Song +6

cs.CR2022

Deduplicating Training Data Mitigates Privacy Risks in Language Models

Nikhil Kandpal, Eric Wallace, Colin Raffel

cs.LG2019

MixMatch: A Holistic Approach to Semi-Supervised Learning

David Berthelot, Nicholas Carlini, Ian Goodfellow +3

stat.ML2020

Top-k Training of GANs: Improving GAN Performance by Throwing Away Bad Samples

Samarth Sinha, Zhengli Zhao, Anirudh Goyal +2

cs.LG2018

Understanding and Improving Interpolation in Autoencoders via an Adversarial Regularizer

David Berthelot, Colin Raffel, Aurko Roy +1

cs.CL2023

Efficient Methods for Natural Language Processing: A Survey

Marcos Treviso, Ji-Ung Lee, Tianchu Ji +19

cs.LG2023

ColD Fusion: Collaborative Descent for Distributed Multitask Finetuning

Shachar Don-Yehiya, Elad Venezian, Colin Raffel +3

cs.LG2017

Training a Subsampling Mechanism in Expectation

Colin Raffel, Dieterich Lawson

cs.LG2022

PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts

Stephen H. Bach, Victor Sanh, Zheng-Xin Yong +24

cs.CL2020

How Much Knowledge Can You Pack Into the Parameters of a Language Model?

Adam Roberts, Colin Raffel, Noam Shazeer

cs.LG2025

ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization

Prateek Yadav, Leshem Choshen, Colin Raffel +1

cs.LG2025

AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution

Fengyuan Liu, Nikhil Kandpal, Colin Raffel

cs.CL2023

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

BigScience Workshop, :, Teven Le Scao +391

stat.ML2018

Is Generator Conditioning Causally Related to GAN Performance?

Augustus Odena, Jacob Buckman, Catherine Olsson +4

cs.CL2022

ByT5: Towards a token-free future with pre-trained byte-to-byte models

Linting Xue, Aditya Barua, Noah Constant +5

cs.CL2025

Scaling Data-Constrained Language Models

Niklas Muennighoff, Alexander M. Rush, Boaz Barak +6

cs.CL2019

Monotonic Infinite Lookback Attention for Simultaneous Machine Translation

Naveen Arivazhagan, Colin Cherry, Wolfgang Macherey +5

eess.AS2019

Imperceptible, Robust, and Targeted Adversarial Examples for Automatic Speech Recognition

Yao Qin, Nicholas Carlini, Ian Goodfellow +2

cs.LG2023

Improving Few-Shot Generalization by Exploring and Exploiting Auxiliary Data

Alon Albalak, Colin Raffel, William Yang Wang

cs.CL2025

Position: The Most Expensive Part of an LLM should be its Training Data

Nikhil Kandpal, Colin Raffel

cs.LG2024

Learning to Route Among Specialized Experts for Zero-Shot Generalization

Mohammed Muqeeth, Haokun Liu, Yufan Liu +1

stat.ML2018

Learning a Latent Space of Multitrack Measures

Ian Simon, Adam Roberts, Colin Raffel +3

cs.LG2025

A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning

Prateek Yadav, Colin Raffel, Mohammed Muqeeth +6

cs.CL2022

What Language Model to Train if You Have One Million GPU Hours?

Teven Le Scao, Thomas Wang, Daniel Hesslow +16

cs.LG2020

Towards GAN Benchmarks Which Require Generalization

Ishaan Gulrajani, Colin Raffel, Luke Metz

cs.LG2026

The Appeal and Reality of Recycling LoRAs with Adaptive Merging

Haokun Liu, Gyung Hyun Je, Marco Ciccone +3

cs.CL2026

TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior

Gül Sena Altıntaş, Malikeh Ehghaghi, Brian Lester +4

cs.LG2026

Model Merging via Data-Free Covariance Estimation

Marawan Gamal Abdel Hameed, Derek Tam, Pascal Jr Tikeng Notsawo +2

cs.LG2022

Merging Models with Fisher-Weighted Averaging

Michael Matena, Colin Raffel

cs.CL2023

Efficient Online Data Mixing For Language Model Pre-Training

Alon Albalak, Liangming Pan, Colin Raffel +1

cs.LG2024

Soft Merging of Experts with Adaptive Routing

Mohammed Muqeeth, Haokun Liu, Colin Raffel

cs.LG2023

Bidirectional Language Models Are Also Few-shot Learners

Ajay Patel, Bryan Li, Mohammad Sadegh Rasooli +3

cs.LG2021

Training Neural Networks with Fixed Sparse Masks

Yi-Lin Sung, Varun Nair, Colin Raffel

cs.CL2021

Improving and Simplifying Pattern Exploiting Training

Derek Tam, Rakesh R Menon, Mohit Bansal +2

cs.CL2024

DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows

Ajay Patel, Colin Raffel, Chris Callison-Burch

cs.CL2023

Evaluating the Factual Consistency of Large Language Models Through News Summarization

Derek Tam, Anisha Mascarenhas, Shiyue Zhang +3

cs.LG2025

The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions

Devin Kwok, Gül Sena Altıntaş, Colin Raffel +1

cs.CL2025

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Guilherme Penedo, Hynek Kydlíček, Vinko Sabolčec +7

cs.CL2025

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

Nikhil Kandpal, Brian Lester, Colin Raffel +24

cs.LG2026

Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models

Malikeh Ehghaghi, Boglárka Ecsedi, Marsha Chechik +1

cs.LG2025

Enhancing Training Data Attribution with Representational Optimization

Weiwei Sun, Haokun Liu, Nikhil Kandpal +2

cs.CL2021

NeurIPS 2020 EfficientQA Competition: Systems, Analyses and Lessons Learned

Sewon Min, Jordan Boyd-Graber, Chris Alberti +50

cs.CL2026

FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale

Ajay Patel, Colin Raffel, Chris Callison-Burch

cs.CL2023

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao +448

cs.AI2017

Learning Hard Alignments with Variational Inference

Dieterich Lawson, Chung-Cheng Chiu, George Tucker +3

cs.LG2022

A Combinatorial Perspective on the Optimization of Shallow ReLU Networks

Michael Matena, Colin Raffel

cs.LG2024

Merging by Matching Models in Task Parameter Subspaces

Derek Tam, Mohit Bansal, Colin Raffel

cs.LG2020

Detecting and Diagnosing Adversarial Images with Class-Conditional Capsule Reconstructions

Yao Qin, Nicholas Frosst, Sara Sabour +3

cs.LG2023

TIES-Merging: Resolving Interference When Merging Models

Prateek Yadav, Derek Tam, Leshem Choshen +2

cs.LG2022

Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$

Adam Roberts, Hyung Won Chung, Anselm Levskaya +40

cs.CL2024

A Survey on Data Selection for Language Models

Alon Albalak, Yanai Elazar, Sang Michael Xie +11

cs.LG2020

FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence

Kihyuk Sohn, David Berthelot, Chun-Liang Li +6