activity
20242026
most citedChain-of-Thought Reasoning In The Wild Is Not Always Faithful

1 citations · 1 across the 2 of their papers we have counts for

collaborators

13 papers

cs.AI2026

Base Models Know How to Reason, Thinking Models Learn When

Constantin Venhoff, Iván Arcuschin, Philip Torr +2

What do thinking language models learn during training that their base models lack? We first present an unsupervised method that discovers a model's reasoning behaviors by training…

cs.AI20261 cited

Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Iván Arcuschin, Jett Janiak, Robert Krzyzanowski +3

Recent studies indicate that when faced with explicit biases in prompts, models often omit mentioning these biases in their Chain-of-Thought (CoT) output, revealing that verbalized…

cs.LG2026

Automatically Finding Reward Model Biases

Atticus Wang, Iván Arcuschin, Arthur Conmy

Reward models are central to large language model (LLM) post-training. However, past work has shown that they can reward spurious or undesirable attributes such as length, format,…

cs.LG2025

Eliciting Secret Knowledge from Language Models

Bartosz Cywiński, Emil Ryd, Rowan Wang +4

We study secret elicitation: discovering knowledge that an AI possesses but does not explicitly verbalize. As a testbed, we train three families of large language models (LLMs) to…

cs.LG2025

Thought Anchors: Which LLM Reasoning Steps Matter?

Paul C. Bogdan, Uzay Macar, Neel Nanda +1

Current frontier large-language models rely on reasoning to achieve state-of-the-art performance. Many existing interpretability are limited in this area, as standard methods have…

cs.LG2025

Understanding Reasoning in Thinking Language Models via Steering Vectors

Constantin Venhoff, Iván Arcuschin, Philip Torr +2

Recent advances in large language models (LLMs) have led to the development of thinking language models that generate extensive internal reasoning chains before producing responses…