1 citations · 1 across the 2 of their papers we have counts for
13 papers
Base Models Know How to Reason, Thinking Models Learn When
Constantin Venhoff, Iván Arcuschin, Philip Torr +2
What do thinking language models learn during training that their base models lack? We first present an unsupervised method that discovers a model's reasoning behaviors by training…
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Iván Arcuschin, Jett Janiak, Robert Krzyzanowski +3
Recent studies indicate that when faced with explicit biases in prompts, models often omit mentioning these biases in their Chain-of-Thought (CoT) output, revealing that verbalized…
Automatically Finding Reward Model Biases
Atticus Wang, Iván Arcuschin, Arthur Conmy
Reward models are central to large language model (LLM) post-training. However, past work has shown that they can reward spurious or undesirable attributes such as length, format,…
Eliciting Secret Knowledge from Language Models
Bartosz CywiÅski, Emil Ryd, Rowan Wang +4
We study secret elicitation: discovering knowledge that an AI possesses but does not explicitly verbalize. As a testbed, we train three families of large language models (LLMs) to…
Thought Anchors: Which LLM Reasoning Steps Matter?
Paul C. Bogdan, Uzay Macar, Neel Nanda +1
Current frontier large-language models rely on reasoning to achieve state-of-the-art performance. Many existing interpretability are limited in this area, as standard methods have…
Understanding Reasoning in Thinking Language Models via Steering Vectors
Constantin Venhoff, Iván Arcuschin, Philip Torr +2
Recent advances in large language models (LLMs) have led to the development of thinking language models that generate extensive internal reasoning chains before producing responses…