6 papers
Solving adversarial examples requires solving exponential misalignment
Alessandro Salvatore, Stanislav Fort, Surya Ganguli
Adversarial attacks - input perturbations imperceptible to humans that fool neural networks - remain both a persistent failure mode in machine learning, and a phenomenon with myste…
Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks
Daniel Kunin, Giovanni Luca Marchetti, Feng Chen +5
What features neural networks learn, and how, remains an open question. In this paper, we introduce Alternating Gradient Flows (AGF), an algorithmic framework that describes the dy…
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
Feng Chen, Allan Raventos, Nan Cheng +2
Recent progress in large language models (LLMs) highlights the power of scaling test-time compute to achieve strong performance on complex tasks, such as mathematical reasoning and…
Features are fate: a theory of transfer learning in high-dimensional regression
Javan Tahir, Surya Ganguli, Grant M. Rotskoff
With the emergence of large-scale pre-trained neural networks, methods to adapt such "foundation" models to data-limited downstream tasks have become a necessity. Fine-tuning, pref…
An analytic theory of creativity in convolutional diffusion models
Mason Kamb, Surya Ganguli
We obtain an analytic, interpretable and predictive theory of creativity in convolutional diffusion models. Indeed, score-matching diffusion models can generate highly original ima…
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
Daniel Kunin, Allan Raventós, Clémentine Dominé +4
While the impressive performance of modern neural networks is often attributed to their capacity to efficiently extract task-relevant features from data, the mechanisms underlying…