102 citations · 154 across the 18 of their papers we have counts for
Showing 2025 · cs.AIShow all
2 papers · 2 filters
cs.AI2025
What Does it Mean for a Neural Network to Learn a "World Model"?
Kenneth Li, Fernanda Viégas, Martin Wattenberg
We propose a set of precise criteria for saying a neural net learns and uses a "world model." The goal is to give an operational meaning to terms that are often used informally, in…
cs.AI2025
The Geometry of Self-Verification in a Task-Specific Reasoning Model
Andrew Lee, Lihao Sun, Chris Wendler +2
How do reasoning models verify their own answers? We study this question by training a model using DeepSeek R1's recipe on the CountDown task. We leverage the fact that preference…