activity
20242026
collaborators

7 papers

cs.LG2026

Structural Inference: Interpreting Small Language Models with Susceptibilities

Garrett Baker, George Wang, Jesse Hoogland +1

We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution,…

stat.ML2025

Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory

Einar Urdshals, Edmund Lau, Jesse Hoogland +2

We study neural network compressibility by using singular learning theory to extend the minimum description length (MDL) principle to singular models like neural networks. Through…

cs.LG2025

Loss Landscape Degeneracy and Stagewise Development in Transformers

Jesse Hoogland, George Wang, Matthew Farrugia-Roberts +3

Deep learning involves navigating a high-dimensional loss landscape over the neural network parameter space. Over the course of training, complex computational structures form and…

cs.LG2025

You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation

Simon Pepin Lehalleur, Jesse Hoogland, Matthew Farrugia-Roberts +5

In this position paper, we argue that understanding the relation between structure in the data distribution and structure in trained models is central to AI alignment. First, we di…

cs.LG2025

Dynamics of Transient Structure in In-Context Linear Regression Transformers

Liam Carroll, Jesse Hoogland, Matthew Farrugia-Roberts +1

Modern deep neural networks display striking examples of rich internal computational structure. Uncovering principles governing the development of such structure is a priority for…

cs.LG2025

Open Problems in Mechanistic Interpretability

Lee Sharkey, Bilal Chughtai, Joshua Batson +26

Mechanistic interpretability aims to understand the computational mechanisms underlying neural networks' capabilities in order to accomplish concrete scientific and engineering goa…