activity
20242026
most citedAgents of Chaos

6 citations · 7 across the 4 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG20261 cited

Two Stages of Folding: Convergent Mechanisms in AI Protein Folding Trunks

Kevin Lu, Jannik Brinkmann, Stefan Huber +4

How do protein structure prediction models fold proteins? We investigate this question through causal interventions on the folding trunks of ESMFold, OpenFold, and Boltz-1. Across…

cs.LG2025

The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis

Aaron Mueller, Jannik Brinkmann, Millicent Li +10

Interpretability provides a toolset for understanding how and why neural networks behave in certain ways. However, there is little unity in the field: most studies employ ad-hoc ev…

cs.LG2025

MIB: A Mechanistic Interpretability Benchmark

Aaron Mueller, Atticus Geiger, Sarah Wiegreffe +20

How can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretabili…

cs.LG2025

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Samuel Marks, Can Rager, Eric J. Michaud +3

We introduce methods for discovering and applying sparse feature circuits. These are causally implicated subnetworks of human-interpretable features for explaining language model b…

cs.LG2025

Position-aware Automatic Circuit Discovery

Tal Haklay, Hadas Orgad, David Bau +2

A widely used strategy to discover and understand language model mechanisms is circuit analysis. A circuit is a minimal subgraph of a model's computation graph that executes a spec…

cs.LG2024

Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models

Adam Karvonen, Benjamin Wright, Can Rager +6

What latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representati…