4 papers
Lost in Aggregation: The Causal Interpretation of the IV Estimand
Danielle Tsao, Krikamol Muandet, Frederick Eberhardt +1
Instrumental variable estimation has emerged as a standard approach to mitigating confounding bias in the social sciences and epidemiology, where conducting randomized experiments…
When the Coffee Feature Activates on Coffins: An Analysis of Feature Extraction and Steering for Mechanistic Interpretability
Raphael Ronge, Markus Maier, Frederick Eberhardt
Recent work by Anthropic on Mechanistic interpretability claims to understand and control Large Language Models by extracting human-interpretable features from their neural activat…
Modeling Discrimination with Causal Abstraction
Milan Mossé, Kara Schechtman, Frederick Eberhardt +1
A person is directly racially discriminated against only if her race caused her worse treatment. This implies that race is an attribute sufficiently separable from other attributes…
Lower Bounds on the Size of Markov Equivalence Classes
Erik Jahn, Frederick Eberhardt, Leonard J. Schulman
Causal discovery algorithms typically recover causal graphs only up to their Markov equivalence classes unless additional parametric assumptions are made. The sizes of these equiva…