4 papers
Exemplar Partitioning for Mechanistic Interpretability
Jessica Rumbelow
We introduce Exemplar Partitioning (EP), an unsupervised method for constructing interpretable feature dictionaries from large language model activations with few…
Explaining Surface Layer Theory Departures in Marine Flux Profiles with Data-Driven Discovery
Jack Foxabbott, Leo Mckee-Reid, Andrew Cusick +8
Monin--Obukhov Similarity Theory (MOST), which underpins nearly all bulk estimates of surface fluxes in the atmospheric surface layer, assumes monotonic wind profiles and verticall…
Benchmarking the Discovery Engine
Jack Foxabbott, Arush Tagade, Andrew Cusick +6
The Discovery Engine is a general purpose automated system for scientific discovery, which combines machine learning with state-of-the-art ML interpretability to enable rapid and r…
Open Problems in Mechanistic Interpretability
Lee Sharkey, Bilal Chughtai, Joshua Batson +26
Mechanistic interpretability aims to understand the computational mechanisms underlying neural networks' capabilities in order to accomplish concrete scientific and engineering goa…