1 citations · 1 across the 1 of their papers we have counts for
7 papers
How Causal Abstraction Underpins Computational Explanation
Atticus Geiger, Jacqueline Harding, Thomas Icard
Explanations of cognitive behavior often appeal to computations over representations. What does it take for a system to implement a given computation over suitable representational…
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
Li Puyin, Jiyuan Tan, Ahmad Jabbar +2
We present a method for diagnosing interpretation in neural networks by identifying an input subspace where a proposed interpretation is highly faithful. Our method is particularly…
HyperSteer: Activation Steering at Scale with Hypernetworks
Jiuding Sun, Sidharth Baskaran, Zhengxuan Wu +3
Steering language models (LMs) by modifying internal activations is a popular approach for controlling text generation. Unsupervised dictionary learning methods, e.g., sparse autoe…
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
Atticus Geiger, Duligur Ibeling, Amir Zur +8
Causal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications…
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
Jiuding Sun, Jing Huang, Sidharth Baskaran +4
Mechanistic interpretability has made great strides in identifying neural network features (e.g., directions in hidden activation space) that mediate concepts(e.g., the birth year…
Combining Causal Models for More Accurate Abstractions of Neural Networks
Theodora-Mara Pîslar, Sara Magliacane, Atticus Geiger
Mechanistic interpretability aims to reverse engineer neural networks by uncovering which high-level algorithms they implement. Causal abstraction provides a precise notion of when…