Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers
Andrew Nam, Henry Conklin, Yukang Yang +3
We present causal head gating (CHG), a scalable method for interpreting the functional roles of attention heads in transformer models. CHG learns soft gates over heads and assigns…
cs.AI2024
Discrete, compositional, and symbolic representations through attractor dynamics
Andrew Nam, Eric Elmoznino, Nikolay Malkin +3
Symbolic systems are powerful frameworks for modeling cognitive processes as they encapsulate the rules and relationships fundamental to many aspects of human reasoning and behavio…