1 citations · 1 across the 3 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026★ 1 cited
How Causal Abstraction Underpins Computational Explanation
Atticus Geiger, Jacqueline Harding, Thomas Icard
Explanations of cognitive behavior often appeal to computations over representations. What does it take for a system to implement a given computation over suitable representational…
cs.LG2026
Transcoder Adapters for Reasoning-Model Diffing
Nathan Hu, Jake Ward, Thomas Icard +1
While reasoning models are increasingly ubiquitous, the effects of reasoning training on a model's internal mechanisms remain poorly understood. In this work, we introduce transcod…
cs.LG2025
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
Jing Huang, Junyi Tao, Thomas Icard +2
Interpretability research now offers a variety of techniques for identifying abstract internal mechanisms in neural networks. Can such techniques be used to predict how models will…