collaborators

5 papers

cs.LG2026

Quantifying the Effect of Test Set Contamination on Generative Evaluations

Rylan Schaeffer, Joshua Kazdan, Baber Abbasi +8

As frontier AI systems are pretrained on web-scale data, test set contamination has become a critical concern for accurately assessing their capabilities. While research has thorou…

cs.LG2025

Evaluating the Robustness of Chinchilla Compute-Optimal Scaling

Rylan Schaeffer, Noam Levi, Andreas Kirsch +4

Hoffman et al (2022)'s Chinchilla paper introduced the principle of compute-optimal scaling, laying a foundation for future scaling of language models. In the years since, however,…

cs.LG2025

Pretraining Scaling Laws for Generative Evaluations of Language Models

Rylan Schaeffer, Noam Levi, Brando Miranda +1

Neural scaling laws have driven the field's ever-expanding exponential growth in parameters, data and compute. While scaling behaviors for pretraining losses and discriminative ben…

cs.LG2025

Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track

Rylan Schaeffer, Joshua Kazdan, Yegor Denisov-Blanch +11

Science progresses by iteratively advancing and correcting humanity's understanding of the world. In machine learning (ML) research, rapid advancements have led to an explosion of…

cs.AI2025

Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization

Willy Chan, Michael Souliman, Jakob Nordhagen +3

Autoformalization, the process of transforming informal mathematical language into formal specifications and proofs remains a difficult task for state-of-the-art (large) language m…