9 papers
Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking
Ali Saheb Pasand, Elvis Dohmatob
Grokking is the phenomenon whereby, unlike the training performance, which peaks early in the training process, the test/generalization performance of a model stagnates over arbitr…
Efficient Refusal Ablation in LLM through Optimal Transport
Geraldin Nanfack, Eugene Belilovsky, Elvis Dohmatob
Safety-aligned language models refuse harmful requests through learned refusal behaviors encoded in their internal representations. Recent activation-based jailbreaking methods cir…
Why Less is More (Sometimes): A Theory of Data Curation
Elvis Dohmatob, Mohammad Pezeshki, Reyhane Askari-Hemmat
This paper introduces a theoretical framework to resolve a central paradox in modern machine learning: When is it better to use less data? This question has become critical as clas…
auto-fpt: Automating Free Probability Theory Calculations for Machine Learning Theory
Arjun Subramonian, Elvis Dohmatob
A large part of modern machine learning theory often involves computing the high-dimensional expected trace of a rational expression of large rectangular random matrices. To symbol…
An Effective Theory of Bias Amplification
Arjun Subramonian, Samuel J. Bell, Levent Sagun +1
Machine learning models can capture and amplify biases present in data, leading to disparate test performance across social groups. To better understand, evaluate, and mitigate the…
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
Reyhane Askari-Hemmat, Mohammad Pezeshki, Elvis Dohmatob +6
Inspired by the principle of deliberate practice in human learning, we propose Deliberate Practice for Synthetic Data Generation (DP), a novel framework that improves sample effici…