2 papers
cs.LG2026
Interpretability Can Be Actionable
Hadas Orgad, Fazl Barez, Tal Haklay +9
Interpretability aims to explain the behavior of deep neural networks. Despite rapid growth, there is mounting concern that much of this work has not translated into practical impa…
cs.LG2025
The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail
Cassandra Goldberg, Chaehyeon Kim, Adam Stein +1
Concept vectors aim to enhance model interpretability by linking internal representations with human-understandable semantics, but their practical utility is often limited by noisy…