collaborators

6 papers

stat.ML2026

There Will Be a Scientific Theory of Deep Learning

Jamie Simon, Daniel Kunin, Alexander Atanasov +11

In this paper, we make the case that a scientific theory of deep learning is emerging. By this we mean a theory which characterizes important properties and statistics of the train…

cs.LG2025

On the creation of narrow AI: hierarchy and nonlocality of neural network skills

Eric J. Michaud, Asher Parker-Sartori, Max Tegmark

We study the problem of creating strong, yet narrow, AI systems. While recent AI progress has been driven by the training of large general-purpose foundation models, the creation o…

cs.LG2025

Efficient Dictionary Learning with Switch Sparse Autoencoders

Anish Mudide, Joshua Engels, Eric J. Michaud +2

Sparse autoencoders (SAEs) are a recent technique for decomposing neural network activations into human-interpretable features. However, in order for SAEs to identify all features…

q-bio.NC2025

The Geometry of Concepts: Sparse Autoencoder Feature Structure

Yuxiao Li, Eric J. Michaud, David D. Baek +3

Sparse autoencoders have recently produced dictionaries of high-dimensional vectors corresponding to the universe of concepts represented by large language models. We find that thi…

cs.LG2025

Not All Language Model Features Are One-Dimensionally Linear

Joshua Engels, Eric J. Michaud, Isaac Liao +2

Recent work has proposed that language models perform computation by manipulating one-dimensional representations of concepts ("features") in activation space. In contrast, we expl…

cs.LG2025

Physics of Skill Learning

Ziming Liu, Yizhou Liu, Eric J. Michaud +2

We aim to understand physics of skill learning, i.e., how skills are learned in neural networks during training. We start by observing the Domino effect, i.e., skills are learned s…