most citedReasoning Models Don't Always Say What They Think

9 citations · 13 across the 5 of their papers we have counts for

collaborators

6 papers

cs.LG2026

Excess Description Length of Learning Generalizable Predictors

Elizabeth Donoway, Hailey Joren, Fabien Roger +1

Understanding whether fine-tuning elicits latent capabilities or teaches new ones is a fundamental question for language model evaluation and safety. We develop a formal informatio…

cs.LG2025

Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment

Nevan Wichers, Aram Ebtekar, Ariana Azarbal +8

Large language models are sometimes trained with imperfect oversight signals, leading to undesired behaviors such as reward hacking and sycophancy. Improving oversight quality can…

cs.CL2025

All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language

Shiyuan Guo, Henry Sleight, Fabien Roger

Detecting harmful AI actions is important as AI agents gain adoption. Chain-of-thought (CoT) monitoring is one method widely used to detect adversarial attacks and AI misalignment.…

cs.CL20259 cited

Reasoning Models Don't Always Say What They Think

Yanda Chen, Joe Benton, Ansh Radhakrishnan +12

Chain-of-thought (CoT) offers a potential boon for AI safety as it allows monitoring a model's CoT to try to understand its intentions and reasoning processes. However, the effecti…

cs.AI20253 cited

Auditing language models for hidden objectives

Samuel Marks, Johannes Treutlein, Trenton Bricken +32

We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objecti…

cs.AI20251 cited

A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management

Simeon Campos, Henry Papadatos, Fabien Roger +3

The recent development of powerful AI systems has highlighted the need for robust risk management frameworks in the AI industry. Although companies have begun to implement safety f…