1 citations · 1 across the 1 of their papers we have counts for
10 papers
Two Stages of Folding: Convergent Mechanisms in AI Protein Folding Trunks
Kevin Lu, Jannik Brinkmann, Stefan Huber +4
How do protein structure prediction models fold proteins? We investigate this question through causal interventions on the folding trunks of ESMFold, OpenFold, and Boltz-1. Across…
Language Models use Lookbacks to Track Beliefs
Nikhil Prakash, Natalie Shapira, Arnab Sen Sharma +5
How do language models (LMs) represent characters' beliefs, especially when those beliefs may differ from reality? This question lies at the heart of understanding the Theory of Mi…
Agents of Chaos
Natalie Shapira, Chris Wendler, Avery Yen +35
We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord acc…
In-Context Learning Without Copying
Kerem Sahin, Sheridan Feucht, Adam Belfki +4
Induction heads are attention heads that perform inductive copying by matching patterns from earlier context and copying their continuations verbatim. As models develop induction h…
The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
Aaron Mueller, Jannik Brinkmann, Millicent Li +10
Interpretability provides a toolset for understanding how and why neural networks behave in certain ways. However, there is little unity in the field: most studies employ ad-hoc ev…
Erasing Conceptual Knowledge from Language Models
Rohit Gandikota, Sheridan Feucht, Samuel Marks +1
In this work, we introduce Erasure of Language Memory (ELM), a principled approach to concept-level unlearning that operates by matching distributions defined by the model's own in…