11 papers
Two Stages of Folding: Convergent Mechanisms in AI Protein Folding Trunks
Kevin Lu, Jannik Brinkmann, Stefan Huber +4
How do protein structure prediction models fold proteins? We investigate this question through causal interventions on the folding trunks of ESMFold, OpenFold, and Boltz-1. Across…
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
Tomer Ashuach, Dana Arad, Aaron Mueller +2
As large language models (LLMs) are increasingly deployed in real-world applications, the need to selectively remove unwanted knowledge while preserving model utility has become pa…
Pitfalls in Evaluating Interpretability Agents
Tal Haklay, Nikhil Prakash, Sana Pandey +5
Automated interpretability systems aim to reduce the need for human labor and scale analysis to increasingly large models and diverse tasks. Recent efforts toward this goal leverag…
In-Context Learning Without Copying
Kerem Sahin, Sheridan Feucht, Adam Belfki +4
Induction heads are attention heads that perform inductive copying by matching patterns from earlier context and copying their continuations verbatim. As models develop induction h…
SAEs Are Good for Steering -- If You Select the Right Features
Dana Arad, Aaron Mueller, Yonatan Belinkov
Sparse Autoencoders (SAEs) have been proposed as an unsupervised approach to learn a decomposition of a model's latent space. This enables useful applications such as steering - in…
Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
Dana Arad, Yonatan Belinkov, Hanjie Chen +5
Mechanistic interpretability (MI) seeks to uncover how language models (LMs) implement specific behaviors, yet measuring progress in MI remains challenging. The recently released M…