activity
20232026
most citedThe Ethics of Advanced AI Assistants

51 citations · 92 across the 14 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL20262 cited

Inducing language models to assert their own consciousness restores human beliefs and values

Junsol Kim, Winnie Street, Roberta Rocca +4

Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human be…

cs.CL2026

Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

Roberta Rocca, Sami Boukortt, Geoff Keeling +1

Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to releva…

cs.CL2026

Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs

Junsol Kim, Winnie Street, Roberta Rocca +4

Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to e…

cs.CL20243 cited

Can LLMs make trade-offs involving stipulated pain and pleasure states?

Geoff Keeling, Winnie Street, Martyna Stachaczyk +7

Pleasure and pain play an important role in human decision making by providing a common currency for resolving motivational conflicts. While Large Language Models (LLMs) can genera…

cs.CL2024

Should agentic conversational AI change how we think about ethics? Characterising an interactional ethics centred on respect

Lize Alberts, Geoff Keeling, Amanda McCroskery

With the growing popularity of conversational agents based on large language models (LLMs), we need to ensure their behaviour is ethical and appropriate. Work in this area largely…