4 papers
Elicitation without Backpropagation: Steering Model Behavior by Optimizing the Latent Posterior
Garrett Baker, Vinayak Pathak, Daniel Murfet +1
In the \emph{latent posterior model} of transformer behavior, the next-token distribution arises from a posterior over latent predictive models conditioned on the context, mixed to…
Structural Inference: Interpreting Small Language Models with Susceptibilities
Garrett Baker, George Wang, Jesse Hoogland +1
We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution,…
Towards Spectroscopy: Susceptibility Clusters in Language Models
Andrew Gordon, Garrett Baker, George Wang +3
Spectroscopy infers the internal structure of physical systems by measuring their response to perturbations. We apply this principle to neural networks: perturbing the data distrib…
Embryology of a Language Model
George Wang, Garrett Baker, Andrew Gordon +1
Understanding how language models develop their internal computational structure is a central problem in the science of deep learning. While susceptibilities, drawn from statistica…