activity
20242026
collaborators
Showing cs.LGShow all

19 papers · 1 filter

cs.LG2026

Internal Data Repetition Destroys Language Models

Jessica Chudnovsky, Joshua Kazdan, Noam Levi +6

Language models are running out of high-quality training data, and even aggressively deduplicated corpora retain some amount of repetition. Earlier controlled studies predated Chin…

cs.LG2026

Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation

Sang Truong, Yuheng Tu, Rylan Schaeffer +1

Scaling laws provide a fundamental framework for understanding the performance of Language Models (LMs), yet deriving them requires prohibitively expensive evaluations across thous…

cs.LG2026

Consensus is Not Verification: Why Crowd Wisdom Strategies Fail for LLM Truthfulness

Yegor Denisov-Blanch, Joshua Kazdan, Jessica Chudnovsky +4

Pass@k and other methods of scaling inference compute can improve language model performance in domains with external verifiers, including mathematics and code, where incorrect can…

cs.LG2026

Scale Dependent Data Duplication

Joshua Kazdan, Noam Levi, Rylan Schaeffer +6

Data duplication during pretraining can degrade generalization and lead to memorization, motivating aggressive deduplication pipelines. However, at web scale, it is unclear what co…

cs.LG2026

Quantifying the Effect of Test Set Contamination on Generative Evaluations

Rylan Schaeffer, Joshua Kazdan, Baber Abbasi +8

As frontier AI systems are pretrained on web-scale data, test set contamination has become a critical concern for accurately assessing their capabilities. While research has thorou…

cs.LG2026

Pretraining Scaling Laws for Generative Evaluations of Language Models

Rylan Schaeffer, Noam Levi, Brando Miranda +1

Neural scaling laws have driven the field's ever-expanding exponential growth in parameters, data and compute. While scaling behaviors for pretraining losses and discriminative ben…