2 papers
cs.LG2026
The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail
Cassandra Goldberg, Chaehyeon Kim, Adam Stein +1
Concept vectors aim to enhance model interpretability by linking internal representations with human-understandable semantics, but their practical utility is often limited by noisy…
cs.LG2026
Interpretability Can Be Actionable
Hadas Orgad, Fazl Barez, Tal Haklay +9
Interpretability aims to explain the behavior of deep neural networks. Despite rapid growth, there is mounting concern that much of this work has not translated into practical impa…