3 papers
cs.LG2026
Understanding Task Representations in Neural Networks via Bayesian Ablation
Andrew Nam, Declan Campbell, Thomas Griffiths +2
Neural networks are powerful tools for cognitive modeling due to their flexibility and emergent properties. However, interpreting their learned representations remains challenging…
cs.CL2026
Levels of Analysis for Large Language Models
Alexander Y. Ku, Declan Campbell, Xuechunzi Bai +10
Modern artificial intelligence systems, such as large language models, are increasingly powerful but also increasingly hard to understand. Recognizing this problem as analogous to…
cs.AI2025
Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers
Andrew Nam, Henry Conklin, Yukang Yang +3
We present causal head gating (CHG), a scalable method for interpreting the functional roles of attention heads in transformer models. CHG learns soft gates over heads and assigns…