activity
20212026
most citedInterpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

54 citations · 125 across the 25 of their papers we have counts for

collaborators
Showing 2025 · cs.LGShow all

7 papers · 2 filters

cs.LG2025

Eliciting Secret Knowledge from Language Models

Bartosz Cywiński, Emil Ryd, Rowan Wang +4

We study secret elicitation: discovering knowledge that an AI possesses but does not explicitly verbalize. As a testbed, we train three families of large language models (LLMs) to…

cs.LG2025

Thought Anchors: Which LLM Reasoning Steps Matter?

Paul C. Bogdan, Uzay Macar, Neel Nanda +1

Current frontier large-language models rely on reasoning to achieve state-of-the-art performance. Many existing interpretability are limited in this area, as standard methods have…

cs.LG2025★ 1 cited

Understanding Reasoning in Thinking Language Models via Steering Vectors

Constantin Venhoff, Iván Arcuschin, Philip Torr +2

Recent advances in large language models (LLMs) have led to the development of thinking language models that generate extensive internal reasoning chains before producing responses…

cs.LG2025

Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning

Stepan Shabalin, Ayush Panda, Dmitrii Kharlapenko +3

Sparse autoencoders are a promising new approach for decomposing language model activations for interpretation and control. They have been applied successfully to vision transforme…

cs.LG2025

Scaling sparse feature circuit finding for in-context learning

Dmitrii Kharlapenko, Stepan Shabalin, Fazl Barez +2

Sparse autoencoders (SAEs) are a popular tool for interpreting large language model activations, but their utility in addressing open questions in interpretability remains unclear.…

cs.LG2025

SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Adam Karvonen, Can Rager, Johnny Lin +12

Sparse autoencoders (SAEs) are a popular technique for interpreting language model activations, and there is extensive recent work on improving SAE effectiveness. However, most pri…