1 citations · 1 across the 2 of their papers we have counts for
4 papers · 1 filter
How Causal Abstraction Underpins Computational Explanation
Atticus Geiger, Jacqueline Harding, Thomas Icard
Explanations of cognitive behavior often appeal to computations over representations. What does it take for a system to implement a given computation over suitable representational…
Combining Causal Models for More Accurate Abstractions of Neural Networks
Theodora-Mara Pîslar, Sara Magliacane, Atticus Geiger
Mechanistic interpretability aims to reverse engineer neural networks by uncovering which high-level algorithms they implement. Causal abstraction provides a precise notion of when…
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
Maheep Chaudhary, Atticus Geiger
A popular new method in mechanistic interpretability is to train high-dimensional sparse autoencoders (SAEs) on neuron activations and use SAE features as the atomic units of analy…
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
Róbert Csordás, Christopher Potts, Christopher D. Manning +1
The Linear Representation Hypothesis (LRH) states that neural networks learn to encode concepts as directions in activation space, and a strong version of the LRH states that model…