16 papers
Elicitation without Backpropagation: Steering Model Behavior by Optimizing the Latent Posterior
Garrett Baker, Vinayak Pathak, Daniel Murfet +1
In the \emph{latent posterior model} of transformer behavior, the next-token distribution arises from a posterior over latent predictive models conditioned on the context, mixed to…
Interpreting Reinforcement Learning Agents with Susceptibilities
Chris Elliott, Einar Urdshals, David Quarel +1
Susceptibilities are a technique for neural network interpretability that studies the response of posterior expectation values of observables to perturbations of the loss. We gener…
Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning
Chris Elliott, Daniel Murfet
These notes introduce the theory of susceptibilities as developed in [arXiv:2504.18274, arXiv:2601.12703] for interpreting neural networks. The susceptibility of an observable …
Linear Response Estimators for Singular Statistical Models
Chris Elliott, Daniel Murfet
We define susceptibilities as a measure of the response of an observable quantity of a parameterized statistical model to a perturbation of the data for a general class of observab…
Structural Inference: Interpreting Small Language Models with Susceptibilities
Garrett Baker, George Wang, Jesse Hoogland +1
We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution,…
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
Chris Elliott, Einar Urdshals, David Quarel +2
Singular learning theory characterizes Bayesian learning as an evolving tradeoff between accuracy and complexity, with transitions between qualitatively different solutions as samp…