4 papers
Estimating the Probability of Sampling a Trained Neural Network at Random
Adam Scherlis, Nora Belrose
We present and analyze an algorithm for estimating the size, under a Gaussian or uniform measure, of a localized neighborhood in neural network parameter space with behavior simila…
Polysemanticity and Capacity in Neural Networks
Adam Scherlis, Kshitij Sachan, Adam S. Jermyn +2
Individual neurons in neural networks often represent a mixture of unrelated features. This phenomenon, called polysemanticity, can make interpreting neural networks more difficult…
Refusal in LLMs is an Affine Function
Thomas Marshall, Adam Scherlis, Nora Belrose
We propose affine concept editing (ACE) as an approach for steering language models' behavior by intervening directly in activations. We begin with an affine decomposition of model…
Understanding Gradient Descent through the Training Jacobian
Nora Belrose, Adam Scherlis
We examine the geometry of neural network training using the Jacobian of trained network parameters with respect to their initial values. Our analysis reveals low-dimensional struc…