2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.LG2024
Understanding Gradient Descent through the Training Jacobian
Nora Belrose, Adam Scherlis
We examine the geometry of neural network training using the Jacobian of trained network parameters with respect to their initial values. Our analysis reveals low-dimensional struc…
cs.LG2024
Refusal in LLMs is an Affine Function
Thomas Marshall, Adam Scherlis, Nora Belrose
We propose affine concept editing (ACE) as an approach for steering language models' behavior by intervening directly in activations. We begin with an affine decomposition of model…
cs.LG2024★ 2 cited
Does Transformer Interpretability Transfer to RNNs?
Gonçalo Paulo, Thomas Marshall, Nora Belrose
Recent advances in recurrent neural network architectures, such as Mamba and RWKV, have enabled RNNs to match or exceed the performance of equal-size transformers in terms of langu…