activity
20212026
most citedInterpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

54 citations · 82 across the 22 of their papers we have counts for

collaborators
Showing cs.LGShow all

19 papers · 1 filter

cs.LG2026

Automatically Finding Reward Model Biases

Atticus Wang, Iván Arcuschin, Arthur Conmy

Reward models are central to large language model (LLM) post-training. However, past work has shown that they can reward spurious or undesirable attributes such as length, format,…

cs.LG2025

Eliciting Secret Knowledge from Language Models

Bartosz Cywiński, Emil Ryd, Rowan Wang +4

We study secret elicitation: discovering knowledge that an AI possesses but does not explicitly verbalize. As a testbed, we train three families of large language models (LLMs) to…

cs.LG2025

Thought Anchors: Which LLM Reasoning Steps Matter?

Paul C. Bogdan, Uzay Macar, Neel Nanda +1

Current frontier large-language models rely on reasoning to achieve state-of-the-art performance. Many existing interpretability are limited in this area, as standard methods have…

cs.LG20251 cited

Understanding Reasoning in Thinking Language Models via Steering Vectors

Constantin Venhoff, Iván Arcuschin, Philip Torr +2

Recent advances in large language models (LLMs) have led to the development of thinking language models that generate extensive internal reasoning chains before producing responses…

cs.LG2025

Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning

Stepan Shabalin, Ayush Panda, Dmitrii Kharlapenko +3

Sparse autoencoders are a promising new approach for decomposing language model activations for interpretation and control. They have been applied successfully to vision transforme…

cs.LG2025

Scaling sparse feature circuit finding for in-context learning

Dmitrii Kharlapenko, Stepan Shabalin, Fazl Barez +2

Sparse autoencoders (SAEs) are a popular tool for interpreting large language model activations, but their utility in addressing open questions in interpretability remains unclear.…