activity
20182026
most citedLinear Representations of Sentiment in Large Language Models

8 citations · 17 across the 14 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG20251 cited

How Causal Abstraction Underpins Computational Explanation

Atticus Geiger, Jacqueline Harding, Thomas Icard

Explanations of cognitive behavior often appeal to computations over representations. What does it take for a system to implement a given computation over suitable representational…

cs.LG2025

Combining Causal Models for More Accurate Abstractions of Neural Networks

Theodora-Mara Pîslar, Sara Magliacane, Atticus Geiger

Mechanistic interpretability aims to reverse engineer neural networks by uncovering which high-level algorithms they implement. Causal abstraction provides a precise notion of when…

cs.LG2024

Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small

Maheep Chaudhary, Atticus Geiger

A popular new method in mechanistic interpretability is to train high-dimensional sparse autoencoders (SAEs) on neuron activations and use SAE features as the atomic units of analy…

cs.LG20241 cited

Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations

Róbert Csordás, Christopher Potts, Christopher D. Manning +1

The Linear Representation Hypothesis (LRH) states that neural networks learn to encode concepts as directions in activation space, and a strong version of the LRH states that model…

cs.LG20241 cited

pyvene: A Library for Understanding and Improving PyTorch Models via Interventions

Zhengxuan Wu, Atticus Geiger, Aryaman Arora +5

Interventions on model-internal states are fundamental operations in many areas of AI, including model editing, steering, robustness, and interpretability. To facilitate such resea…

cs.LG2024

A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments

Zhengxuan Wu, Atticus Geiger, Jing Huang +4

We respond to the recent paper by Makelov et al. (2023), which reviews subspace interchange intervention methods like distributed alignment search (DAS; Geiger et al. 2023) and cla…