4 citations · 4 across the 3 of their papers we have counts for
11 papers
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
Aaron Mueller, Andrew Lee, Shruti Joshi +3
A goal of interpretability is to recover disentangled representations of latent concepts (features) from the activations of neural networks. The quality of features is typically ev…
Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions
Andrew Lee, Fernanda Viégas, Martin Wattenberg
While researchers are finding concepts represented as linear directions in language models, a bag of linear directions fails to capture relational structure. To better understand t…
Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
Thomas Fel, Binxu Wang, Michael A. Lepori +8
DINOv2 is routinely deployed to recognize objects, scenes, and actions; yet the nature of what it perceives remains unknown. As a working baseline, we adopt the Linear Representati…
Decomposing Query-Key Feature Interactions Using Contrastive Covariances
Andrew Lee, Yonatan Belinkov, Fernanda Viégas +1
Despite the central role of attention heads in Transformers, we lack tools to understand why a model attends to a particular token. To address this, we study the query-key (QK) spa…
Better World Models Can Lead to Better Post-Training Performance
Prakhar Gupta, Henry Conklin, Sarah-Jane Leslie +1
In this work we study how explicit world-modeling objectives affect the internal representations and downstream capability of Transformers across different training stages. We use…
Why Can't Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls
Xiaoyan Bai, Itamar Pres, Yuntian Deng +5
Language models are increasingly capable, yet still fail at a seemingly simple task of multi-digit multiplication. In this work, we study why, by reverse-engineering a model that s…