7 papers
Structural Inference: Interpreting Small Language Models with Susceptibilities
Garrett Baker, George Wang, Jesse Hoogland +1
We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution,…
Patterning: The Dual of Interpretability
George Wang, Daniel Murfet
Mechanistic interpretability aims to understand how neural networks generalize beyond their training data by reverse-engineering their internal structures. We introduce patterning…
Towards Spectroscopy: Susceptibility Clusters in Language Models
Andrew Gordon, Garrett Baker, George Wang +3
Spectroscopy infers the internal structure of physical systems by measuring their response to perturbations. We apply this principle to neural networks: perturbing the data distrib…
Embryology of a Language Model
George Wang, Garrett Baker, Andrew Gordon +1
Understanding how language models develop their internal computational structure is a central problem in the science of deep learning. While susceptibilities, drawn from statistica…
Loss Landscape Degeneracy and Stagewise Development in Transformers
Jesse Hoogland, George Wang, Matthew Farrugia-Roberts +3
Deep learning involves navigating a high-dimensional loss landscape over the neural network parameter space. Over the course of training, complex computational structures form and…
You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation
Simon Pepin Lehalleur, Jesse Hoogland, Matthew Farrugia-Roberts +5
In this position paper, we argue that understanding the relation between structure in the data distribution and structure in trained models is central to AI alignment. First, we di…