54 citations · 82 across the 22 of their papers we have counts for
19 papers · 1 filter
Automatically Finding Reward Model Biases
Atticus Wang, Iván Arcuschin, Arthur Conmy
Reward models are central to large language model (LLM) post-training. However, past work has shown that they can reward spurious or undesirable attributes such as length, format,…
Eliciting Secret Knowledge from Language Models
Bartosz Cywiński, Emil Ryd, Rowan Wang +4
We study secret elicitation: discovering knowledge that an AI possesses but does not explicitly verbalize. As a testbed, we train three families of large language models (LLMs) to…
Thought Anchors: Which LLM Reasoning Steps Matter?
Paul C. Bogdan, Uzay Macar, Neel Nanda +1
Current frontier large-language models rely on reasoning to achieve state-of-the-art performance. Many existing interpretability are limited in this area, as standard methods have…
Understanding Reasoning in Thinking Language Models via Steering Vectors
Constantin Venhoff, Iván Arcuschin, Philip Torr +2
Recent advances in large language models (LLMs) have led to the development of thinking language models that generate extensive internal reasoning chains before producing responses…
Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning
Stepan Shabalin, Ayush Panda, Dmitrii Kharlapenko +3
Sparse autoencoders are a promising new approach for decomposing language model activations for interpretation and control. They have been applied successfully to vision transforme…
Scaling sparse feature circuit finding for in-context learning
Dmitrii Kharlapenko, Stepan Shabalin, Fazl Barez +2
Sparse autoencoders (SAEs) are a popular tool for interpreting large language model activations, but their utility in addressing open questions in interpretability remains unclear.…