Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Refusal in LLMs is an Affine Function
Thomas Marshall, Adam Scherlis, Nora Belrose
We propose affine concept editing (ACE) as an approach for steering language models' behavior by intervening directly in activations. We begin with an affine decomposition of model…
cs.LG2024
Does Transformer Interpretability Transfer to RNNs?
Gonçalo Paulo, Thomas Marshall, Nora Belrose
Recent advances in recurrent neural network architectures, such as Mamba and RWKV, have enabled RNNs to match or exceed the performance of equal-size transformers in terms of langu…