activity
20242026
collaborators

9 papers

cs.LG2026

Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking

Ali Saheb Pasand, Elvis Dohmatob

Grokking is the phenomenon whereby, unlike the training performance, which peaks early in the training process, the test/generalization performance of a model stagnates over arbitr…

cs.LG2026

Efficient Refusal Ablation in LLM through Optimal Transport

Geraldin Nanfack, Eugene Belilovsky, Elvis Dohmatob

Safety-aligned language models refuse harmful requests through learned refusal behaviors encoded in their internal representations. Recent activation-based jailbreaking methods cir…

cs.LG2025

Why Less is More (Sometimes): A Theory of Data Curation

Elvis Dohmatob, Mohammad Pezeshki, Reyhane Askari-Hemmat

This paper introduces a theoretical framework to resolve a central paradox in modern machine learning: When is it better to use less data? This question has become critical as clas…

cs.LG2025

auto-fpt: Automating Free Probability Theory Calculations for Machine Learning Theory

Arjun Subramonian, Elvis Dohmatob

A large part of modern machine learning theory often involves computing the high-dimensional expected trace of a rational expression of large rectangular random matrices. To symbol…

cs.LG2025

An Effective Theory of Bias Amplification

Arjun Subramonian, Samuel J. Bell, Levent Sagun +1

Machine learning models can capture and amplify biases present in data, leading to disparate test performance across social groups. To better understand, evaluate, and mitigate the…

cs.LG2025

Improving the Scaling Laws of Synthetic Data with Deliberate Practice

Reyhane Askari-Hemmat, Mohammad Pezeshki, Elvis Dohmatob +6

Inspired by the principle of deliberate practice in human learning, we propose Deliberate Practice for Synthetic Data Generation (DP), a novel framework that improves sample effici…