3 papers
cs.LG2026
Understanding Task Representations in Neural Networks via Bayesian Ablation
Andrew Nam, Declan Campbell, Thomas Griffiths +2
Neural networks are powerful tools for cognitive modeling due to their flexibility and emergent properties. However, interpreting their learned representations remains challenging…
cs.LG2025
Better World Models Can Lead to Better Post-Training Performance
Prakhar Gupta, Henry Conklin, Sarah-Jane Leslie +1
In this work we study how explicit world-modeling objectives affect the internal representations and downstream capability of Transformers across different training stages. We use…
cs.AI2025
Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers
Andrew Nam, Henry Conklin, Yukang Yang +3
We present causal head gating (CHG), a scalable method for interpreting the functional roles of attention heads in transformer models. CHG learns soft gates over heads and assigns…