4 papers · 1 filter
Think Before You Lie: How Reasoning Leads to Honesty
Ann Yuan, Asma Ghandeharioun, Carter Blum +6
While existing evaluations of large language models (LLMs) measure deception rates, the underlying conditions that give rise to deceptive behavior are poorly understood. We investi…
What Does it Mean for a Neural Network to Learn a "World Model"?
Kenneth Li, Fernanda Viégas, Martin Wattenberg
We propose a set of precise criteria for saying a neural net learns and uses a "world model." The goal is to give an operational meaning to terms that are often used informally, in…
The Geometry of Self-Verification in a Task-Specific Reasoning Model
Andrew Lee, Lihao Sun, Chris Wendler +2
How do reasoning models verify their own answers? We study this question by training a model using DeepSeek R1's recipe on the CountDown task. We leverage the fact that preference…
Relational Composition in Neural Networks: A Survey and Call to Action
Martin Wattenberg, Fernanda B. Viégas
Many neural nets appear to represent data as linear combinations of "feature vectors." Algorithms for discovering these vectors have seen impressive recent success. However, we arg…