From the 1 of 5 linked papers with an AI index.
5 papers
Pretraining Data Can Be Poisoned through Computational Propaganda
Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith +2
The paper shows that language model pretraining data can be poisoned through publicly editable web discussion pages, and introduces a method called HalfLife to estimate how much ma…
Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
Rahul Nadkarni, Yanai Elazar, Hila Gonen +1
We present an experimental recipe for studying the relationship between training data and language model (LM) behavior. We outline steps for intervening on data batches -- i.e., ``…
Evaluating -Gram Novelty of Language Models Using Rusty-DAWG
William Merrill, Noah A. Smith, Yanai Elazar
How novel are texts generated by language models (LMs) relative to their training corpora? In this work, we investigate the extent to which modern LMs generate -grams from their…
DataDecide: How to Predict Best Pretraining Data with Small Experiments
Ian Magnusson, Nguyen Tai, Ben Bogin +10
Because large language models are expensive to pretrain on different datasets, using smaller-scale experiments to decide on data is crucial for reducing costs. Which benchmarks and…
On Linear Representations and Pretraining Data Frequency in Language Models
Jack Merullo, Noah A. Smith, Sarah Wiegreffe +1
Pretraining data has a direct impact on the behaviors and quality of language models (LMs), but we only understand the most basic principles of this relationship. While most work f…