1 citations · 1 across the 2 of their papers we have counts for
7 papers
How Causal Abstraction Underpins Computational Explanation
Atticus Geiger, Jacqueline Harding, Thomas Icard
Explanations of cognitive behavior often appeal to computations over representations. What does it take for a system to implement a given computation over suitable representational…
A Communication-First Account of Explanation
Jacqueline Harding, Tobias Gerstenberg, Thomas Icard
This paper develops a formal account of causal explanation, grounded in a theory of conversational pragmatics, and inspired by the interventionist idea that explanation is about as…
Evaluating Commercial AI Chatbots as News Intermediaries
Mirac Suzgun, Emily Shen, Federico Bianchi +5
AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their proprietary search integratio…
Transcoder Adapters for Reasoning-Model Diffing
Nathan Hu, Jake Ward, Thomas Icard +1
While reasoning models are increasingly ubiquitous, the effects of reasoning training on a model's internal mechanisms remain poorly understood. In this work, we introduce transcod…
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
Jing Huang, Junyi Tao, Thomas Icard +2
Interpretability research now offers a variety of techniques for identifying abstract internal mechanisms in neural networks. Can such techniques be used to predict how models will…
Modeling Discrimination with Causal Abstraction
Milan Mossé, Kara Schechtman, Frederick Eberhardt +1
A person is directly racially discriminated against only if her race caused her worse treatment. This implies that race is an attribute sufficiently separable from other attributes…