15 citations · 27 across the 3 of their papers we have counts for
3 papers
Modular addition without black-boxes: Compressing explanations of MLPs that compute numerical integration
Chun Hei Yip, Rajashree Agrawal, Lawrence Chan +1
The goal of mechanistic interpretability is discovering simpler, low-rank algorithms implemented by models. While we can compress activations into features, compressing nonlinear f…
Evaluating Language-Model Agents on Realistic Autonomous Tasks
Megan Kinniment, Lucas Jun Koba Sato, Haoxing Du +10
In this report, we explore the ability of language model agents to acquire resources, create copies of themselves, and adapt to novel challenges they encounter in the wild. We refe…
A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations
Bilal Chughtai, Lawrence Chan, Neel Nanda
Universality is a key hypothesis in mechanistic interpretability -- that different models learn similar features and circuits when trained on similar tasks. In this work, we study…