#theory of mind
6 papers · 1 filter
Inducing language models to assert their own consciousness restores human beliefs and values
Junsol Kim, Winnie Street, Roberta Rocca +4
The paper investigates how safety fine‑tuning of large language models reduces their tendency to attribute consciousness to themselves, animals, and objects, and shows that reversi…
Using Theory of Mind to Arbitrate between Social and Non-social Learning
Lance Ying, Ryan Truong, Joshua B. Tenenbaum +1
The paper proposes a Rational Mentalizing model that uses Theory of Mind to decide when to learn from others versus direct experience, and shows that this model matches human behav…
A Causal Model of Theory of Mind in Conflict for Artificial Intelligence
Nikolos Gurney
The paper proposes a structural causal model that determines when an AI system should engage Theory of Mind reasoning in conflict situations, using a DAG to link situational and ag…
Egocentric Bias in Vision-Language Models
Maijunxian Wang, Yijiang Li, Bingyang Wang +6
The paper introduces FlipSet, a benchmark that tests vision‑language models on Level‑2 visual perspective taking by requiring them to mentally rotate 2D character strings, and find…
Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points
Roberta Rocca, Sami Boukortt, Geoff Keeling +1
The paper proposes a new two‑player dialogue game, the Epistemic Asymmetry Schelling Task (EAST), to assess Theory of Mind and epistemic tracking abilities in large language models…
MindClaw: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
Ruoxuan Zhang, Qiaoqiao Wan, Zhengguang Wang +4
The paper introduces MindClaw, a closed-loop framework that lets embodied agents reason about human mental states in real time and intervene only when assistance is needed.