5 papers
Patterning: The Dual of Interpretability
George Wang, Daniel Murfet
Mechanistic interpretability aims to understand how neural networks generalize beyond their training data by reverse-engineering their internal structures. We introduce patterning…
Towards Spectroscopy: Susceptibility Clusters in Language Models
Andrew Gordon, Garrett Baker, George Wang +3
Spectroscopy infers the internal structure of physical systems by measuring their response to perturbations. We apply this principle to neural networks: perturbing the data distrib…
Embryology of a Language Model
George Wang, Garrett Baker, Andrew Gordon +1
Understanding how language models develop their internal computational structure is a central problem in the science of deep learning. While susceptibilities, drawn from statistica…
You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation
Simon Pepin Lehalleur, Jesse Hoogland, Matthew Farrugia-Roberts +5
In this position paper, we argue that understanding the relation between structure in the data distribution and structure in trained models is central to AI alignment. First, we di…
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
George Wang, Jesse Hoogland, Stan van Wingerden +2
We introduce refined variants of the Local Learning Coefficient (LLC), a measure of model complexity grounded in singular learning theory, to study the development of internal stru…