Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Interpretability Can Be Actionable
Hadas Orgad, Fazl Barez, Tal Haklay +9
Interpretability aims to explain the behavior of deep neural networks. Despite rapid growth, there is mounting concern that much of this work has not translated into practical impa…
cs.LG2025
MIB: A Mechanistic Interpretability Benchmark
Aaron Mueller, Atticus Geiger, Sarah Wiegreffe +20
How can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretabili…
cs.LG2025
Position-aware Automatic Circuit Discovery
Tal Haklay, Hadas Orgad, David Bau +2
A widely used strategy to discover and understand language model mechanisms is circuit analysis. A circuit is a minimal subgraph of a model's computation graph that executes a spec…