10 papers
On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness
Matthew Nguyen, Kyle Cox, Austin Meek +1
Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their reasoning, it is p…
Base Models Know How to Reason, Thinking Models Learn When
Constantin Venhoff, Iván Arcuschin, Philip Torr +2
What do thinking language models learn during training that their base models lack? We first present an unsupervised method that discovers a model's reasoning behaviors by training…
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Iván Arcuschin, Jett Janiak, Robert Krzyzanowski +3
Recent studies indicate that when faced with explicit biases in prompts, models often omit mentioning these biases in their Chain-of-Thought (CoT) output, revealing that verbalized…
Biases in the Blind Spot: Detecting What LLMs Fail to Mention
Iván Arcuschin, David Chanin, Adrià Garriga-Alonso +1
Large Language Models (LLMs) often provide chain-of-thought (CoT) reasoning traces that appear plausible, but may hide internal biases. We call these unverbalized biases. Monitorin…
Automatically Finding Reward Model Biases
Atticus Wang, Iván Arcuschin, Arthur Conmy
Reward models are central to large language model (LLM) post-training. However, past work has shown that they can reward spurious or undesirable attributes such as length, format,…
Mind the Performance Gap: Capability-Behavior Trade-offs in Feature Steering
Eitan Sprejer, Oscar AgustÃn Stanchi, MarÃa Victoria Carro +2
Feature steering has emerged as a promising approach for controlling LLM behavior through direct manipulation of internal representations, offering advantages over prompt engineeri…