9 citations · 9 across the 5 of their papers we have counts for
4 papers · 1 filter
You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation
Simon Pepin Lehalleur, Jesse Hoogland, Matthew Farrugia-Roberts +5
In this position paper, we argue that understanding the relation between structure in the data distribution and structure in trained models is central to AI alignment. First, we di…
Dynamics of Transient Structure in In-Context Linear Regression Transformers
Liam Carroll, Jesse Hoogland, Matthew Farrugia-Roberts +1
Modern deep neural networks display striking examples of rich internal computational structure. Uncovering principles governing the development of such structure is a priority for…
Open Problems in Mechanistic Interpretability
Lee Sharkey, Bilal Chughtai, Joshua Batson +26
Mechanistic interpretability aims to understand the computational mechanisms underlying neural networks' capabilities in order to accomplish concrete scientific and engineering goa…
Dynamical versus Bayesian Phase Transitions in a Toy Model of Superposition
Zhongtian Chen, Edmund Lau, Jake Mendel +2
We investigate phase transitions in a Toy Model of Superposition (TMS) using Singular Learning Theory (SLT). We derive a closed formula for the theoretical loss and, in the case of…