works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.AI2026

Pretraining Data Can Be Poisoned through Computational Propaganda

Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith +2

The paper shows that language model pretraining data can be poisoned through publicly editable web discussion pages, and introduces a method called HalfLife to estimate how much ma…

cs.CL2026

Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior

Rahul Nadkarni, Yanai Elazar, Hila Gonen +1

We present an experimental recipe for studying the relationship between training data and language model (LM) behavior. We outline steps for intervening on data batches -- i.e., ``…

cs.CL2025

Evaluating -Gram Novelty of Language Models Using Rusty-DAWG

William Merrill, Noah A. Smith, Yanai Elazar

How novel are texts generated by language models (LMs) relative to their training corpora? In this work, we investigate the extent to which modern LMs generate -grams from their…

cs.LG2025

DataDecide: How to Predict Best Pretraining Data with Small Experiments

Ian Magnusson, Nguyen Tai, Ben Bogin +10

Because large language models are expensive to pretrain on different datasets, using smaller-scale experiments to decide on data is crucial for reducing costs. Which benchmarks and…

cs.CL2025

On Linear Representations and Pretraining Data Frequency in Language Models

Jack Merullo, Noah A. Smith, Sarah Wiegreffe +1

Pretraining data has a direct impact on the behaviors and quality of language models (LMs), but we only understand the most basic principles of this relationship. While most work f…