consciousness attribution 1language model alignment 1mind perception 1safety fine-tuning 1theory of mind 1
From the 1 of 2 linked papers with an AI index.
2 citations · 2 across the 1 of their papers we have counts for
2 papers
cs.CL2026★ 2 cited
Inducing language models to assert their own consciousness restores human beliefs and values
Junsol Kim, Winnie Street, Roberta Rocca +4
The paper investigates how safety fine‑tuning of large language models reduces their tendency to attribute consciousness to themselves, animals, and objects, and shows that reversi…
cs.CL2026
Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs
Junsol Kim, Winnie Street, Roberta Rocca +4
Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to e…