activity
20242026
collaborators
Showing cs.LGShow all

14 papers · 1 filter

cs.LG2026

Elicitation without Backpropagation: Steering Model Behavior by Optimizing the Latent Posterior

Garrett Baker, Vinayak Pathak, Daniel Murfet +1

In the \emph{latent posterior model} of transformer behavior, the next-token distribution arises from a posterior over latent predictive models conditioned on the context, mixed to…

cs.LG2026

Interpreting Reinforcement Learning Agents with Susceptibilities

Chris Elliott, Einar Urdshals, David Quarel +1

Susceptibilities are a technique for neural network interpretability that studies the response of posterior expectation values of observables to perturbations of the loss. We gener…

cs.LG2026

Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning

Chris Elliott, Daniel Murfet

These notes introduce the theory of susceptibilities as developed in [arXiv:2504.18274, arXiv:2601.12703] for interpreting neural networks. The susceptibility of an observable

cs.LG2026

Structural Inference: Interpreting Small Language Models with Susceptibilities

Garrett Baker, George Wang, Jesse Hoogland +1

We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution,…

cs.LG2026

Stagewise Reinforcement Learning and the Geometry of the Regret Landscape

Chris Elliott, Einar Urdshals, David Quarel +2

Singular learning theory characterizes Bayesian learning as an evolving tradeoff between accuracy and complexity, with transitions between qualitatively different solutions as samp…

cs.LG2026

Patterning: The Dual of Interpretability

George Wang, Daniel Murfet

Mechanistic interpretability aims to understand how neural networks generalize beyond their training data by reverse-engineering their internal structures. We introduce patterning…