From the 2 of 10 linked papers with an AI index.
2 citations · 2 across the 2 of their papers we have counts for
10 papers
Inducing language models to assert their own consciousness restores human beliefs and values
Junsol Kim, Winnie Street, Roberta Rocca +4
The paper investigates how safety fine‑tuning of large language models reduces their tendency to attribute consciousness to themselves, animals, and objects, and shows that reversi…
Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points
Roberta Rocca, Sami Boukortt, Geoff Keeling +1
The paper proposes a new two‑player dialogue game, the Epistemic Asymmetry Schelling Task (EAST), to assess Theory of Mind and epistemic tracking abilities in large language models…
Chuck, Wilson and the emergence of artificial minds in human-AI conversations
Geoff Keeling, Winnie Street
Large Language Models (LLMs) can simulate person-like things which at least appear to have stable behavioural and psychological dispositions. Call these things characters. Are char…
Epistemic Trust as a Mechanism for Ethics Integration: Failure Modes and Design Principles from 70 Moral Imagination Workshops
Benjamin Lange, Geoff Keeling, Kyle Pedersen +4
Bottom-up responsible innovation initiatives seek to empower technology development teams to engage in ethical reflection, yet such interventions frequently fail to achieve practit…
Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs
Junsol Kim, Winnie Street, Roberta Rocca +4
Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to e…
Architecting Trust in Artificial Epistemic Agents
Nahema Marchal, Stephanie Chan, Matija Franklin +5
Large language models increasingly function as epistemic agents -- entities that can 1) autonomously pursue epistemic goals and 2) actively shape our shared knowledge environment.…