works on

From the 2 of 10 linked papers with an AI index.

most citedInducing language models to assert their own consciousness restores human beliefs and values

2 citations · 2 across the 2 of their papers we have counts for

collaborators

10 papers

cs.CL20262 cited

Inducing language models to assert their own consciousness restores human beliefs and values

Junsol Kim, Winnie Street, Roberta Rocca +4

The paper investigates how safety fine‑tuning of large language models reduces their tendency to attribute consciousness to themselves, animals, and objects, and shows that reversi…

cs.CL2026

Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

Roberta Rocca, Sami Boukortt, Geoff Keeling +1

The paper proposes a new two‑player dialogue game, the Epistemic Asymmetry Schelling Task (EAST), to assess Theory of Mind and epistemic tracking abilities in large language models…

cs.HC2026

Chuck, Wilson and the emergence of artificial minds in human-AI conversations

Geoff Keeling, Winnie Street

Large Language Models (LLMs) can simulate person-like things which at least appear to have stable behavioural and psychological dispositions. Call these things characters. Are char…

cs.CY2026

Epistemic Trust as a Mechanism for Ethics Integration: Failure Modes and Design Principles from 70 Moral Imagination Workshops

Benjamin Lange, Geoff Keeling, Kyle Pedersen +4

Bottom-up responsible innovation initiatives seek to empower technology development teams to engage in ethical reflection, yet such interventions frequently fail to achieve practit…

cs.CL2026

Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs

Junsol Kim, Winnie Street, Roberta Rocca +4

Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to e…

cs.AI2026

Architecting Trust in Artificial Epistemic Agents

Nahema Marchal, Stephanie Chan, Matija Franklin +5

Large language models increasingly function as epistemic agents -- entities that can 1) autonomously pursue epistemic goals and 2) actively shape our shared knowledge environment.…