works on

From the 2 of 16 linked papers with an AI index.

activity
20242026
most citedInducing language models to assert their own consciousness restores human beliefs and values

2 citations · 2 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL20262 cited

Inducing language models to assert their own consciousness restores human beliefs and values

Junsol Kim, Winnie Street, Roberta Rocca +4

The paper investigates how safety fine‑tuning of large language models reduces their tendency to attribute consciousness to themselves, animals, and objects, and shows that reversi…

cs.CL2026

Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

Roberta Rocca, Sami Boukortt, Geoff Keeling +1

The paper proposes a new two‑player dialogue game, the Epistemic Asymmetry Schelling Task (EAST), to assess Theory of Mind and epistemic tracking abilities in large language models…

cs.CL2026

Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs

Junsol Kim, Winnie Street, Roberta Rocca +4

Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to e…

cs.CL2024

Can LLMs make trade-offs involving stipulated pain and pleasure states?

Geoff Keeling, Winnie Street, Martyna Stachaczyk +7

Pleasure and pain play an important role in human decision making by providing a common currency for resolving motivational conflicts. While Large Language Models (LLMs) can genera…

cs.CL2024

Should agentic conversational AI change how we think about ethics? Characterising an interactional ethics centred on respect

Lize Alberts, Geoff Keeling, Amanda McCroskery

With the growing popularity of conversational agents based on large language models (LLMs), we need to ensure their behaviour is ethical and appropriate. Work in this area largely…