activity
20242026
collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Patterning in Practice: Debiasing Reward Models with Susceptibilities

George Wang, Elizabeth Donoway, Daniel Murfet

Reward models trained on human preferences are known to suffer from length, formatting, and other stylistic biases. In this paper we use patterning, which reweights each preference…

cs.LG2026

Patterning: The Dual of Interpretability

George Wang, Daniel Murfet

Mechanistic interpretability aims to understand how neural networks generalize beyond their training data by reverse-engineering their internal structures. We introduce patterning…

cs.LG2026

Towards Spectroscopy: Susceptibility Clusters in Language Models

Andrew Gordon, Garrett Baker, George Wang +3

Spectroscopy infers the internal structure of physical systems by measuring their response to perturbations. We apply this principle to neural networks: perturbing the data distrib…

cs.LG2025

Embryology of a Language Model

George Wang, Garrett Baker, Andrew Gordon +1

Understanding how language models develop their internal computational structure is a central problem in the science of deep learning. While susceptibilities, drawn from statistica…

cs.LG2025

Structural Inference: Interpreting Small Language Models with Susceptibilities

Garrett Baker, George Wang, Jesse Hoogland +1

We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution,…

cs.LG2025

You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation

Simon Pepin Lehalleur, Jesse Hoogland, Matthew Farrugia-Roberts +5

In this position paper, we argue that understanding the relation between structure in the data distribution and structure in trained models is central to AI alignment. First, we di…