papers

Publications (19)

cs.CY2024

The Ethics of Advanced AI Assistants

Iason Gabriel, Arianna Manzini, Geoff Keeling +54

This paper focuses on the opportunities and the ethical and societal risks posed by advanced AI assistants. We define advanced AI assistants as artificial agents with natural langu…

cs.CL2026

Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

Roberta Rocca, Sami Boukortt, Geoff Keeling +1

The paper proposes a new two‑player dialogue game, the Epistemic Asymmetry Schelling Task (EAST), to assess Theory of Mind and epistemic tracking abilities in large language models…

#theory of mind#large language models#dialogue games#epistemic reasoning
cs.HC2026

Chuck, Wilson and the emergence of artificial minds in human-AI conversations

Geoff Keeling, Winnie Street

Large Language Models (LLMs) can simulate person-like things which at least appear to have stable behavioural and psychological dispositions. Call these things characters. Are char…

cs.AI2024

LLMs achieve adult human performance on higher-order theory of mind tasks

Winnie Street, John Oliver Siy, Geoff Keeling +7

This paper examines the extent to which large language models (LLMs) have developed higher-order theory of mind (ToM); the human ability to reason about multiple mental and emotion…

cs.CL2024

Should agentic conversational AI change how we think about ethics? Characterising an interactional ethics centred on respect

Lize Alberts, Geoff Keeling, Amanda McCroskery

With the growing popularity of conversational agents based on large language models (LLMs), we need to ensure their behaviour is ethical and appropriate. Work in this area largely…

cs.CL2026

Inducing language models to assert their own consciousness restores human beliefs and values

Junsol Kim, Winnie Street, Roberta Rocca +4

The paper investigates how safety fine‑tuning of large language models reduces their tendency to attribute consciousness to themselves, animals, and objects, and shows that reversi…

#language model alignment#consciousness attribution#mind perception#safety fine-tuning
cs.CY2026

Epistemic Trust as a Mechanism for Ethics Integration: Failure Modes and Design Principles from 70 Moral Imagination Workshops

Benjamin Lange, Geoff Keeling, Kyle Pedersen +4

Bottom-up responsible innovation initiatives seek to empower technology development teams to engage in ethical reflection, yet such interventions frequently fail to achieve practit…

cs.HC2025

We Need Accountability in Human-AI Agent Relationships

Benjamin Lange, Geoff Keeling, Arianna Manzini +1

We argue that accountability mechanisms are needed in human-AI agent relationships to ensure alignment with user and societal interests. We propose a framework according to which A…

cs.CY2023

Engaging Engineering Teams Through Moral Imagination: A Bottom-Up Approach for Responsible Innovation and Ethical Culture Change in Technology Companies

Benjamin Lange, Geoff Keeling, Amanda McCroskery +5

We propose a "Moral Imagination" methodology to facilitate a culture of responsible innovation for engineering and product teams in technology companies. Our approach has been oper…

cs.AI2024

On the attribution of confidence to large language models

Geoff Keeling, Winnie Street

Credences are mental states corresponding to degrees of confidence in propositions. Attribution of credences to Large Language Models (LLMs) is commonplace in the empirical literat…

cs.CY2025

We Need a New Ethics for a World of AI Agents

Iason Gabriel, Geoff Keeling, Arianna Manzini +1

The deployment of capable AI agents raises fresh questions about safety, human-machine relationships and social coordination. We argue for greater engagement by scientists, scholar…

cs.AI2026

Architecting Trust in Artificial Epistemic Agents

Nahema Marchal, Stephanie Chan, Matija Franklin +5

Large language models increasingly function as epistemic agents -- entities that can 1) autonomously pursue epistemic goals and 2) actively shape our shared knowledge environment.…

cs.AI2025

Deflating Deflationism: A Critical Perspective on Debunking Arguments Against LLM Mentality

Alex Grzankowski, Geoff Keeling, Henry Shevlin +1

Many people feel compelled to interpret, describe, and respond to Large Language Models (LLMs) as if they possess inner mental lives similar to our own. Responses to this phenomeno…

cs.CL2026

Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs

Junsol Kim, Winnie Street, Roberta Rocca +4

Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to e…

cs.LG2022

On the Opportunities and Risks of Foundation Models

Rishi Bommasani, Drew A. Hudson, Ehsan Adeli +111

AI is undergoing a paradigm shift with the rise of models (e.g., BERT, DALL-E, GPT-3) that are trained on broad data at scale and are adaptable to a wide range of downstream tasks.…

cs.CY2023

Algorithmic Bias, Generalist Models,and Clinical Medicine

Geoff Keeling

The technical landscape of clinical machine learning is shifting in ways that destabilize pervasive assumptions about the nature and causes of algorithmic bias. On one hand, the do…

cs.CL2024

Can LLMs make trade-offs involving stipulated pain and pleasure states?

Geoff Keeling, Winnie Street, Martyna Stachaczyk +7

Pleasure and pain play an important role in human decision making by providing a common currency for resolving motivational conflicts. While Large Language Models (LLMs) can genera…

cs.CY2024

A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI

Seliem El-Sayed, Canfer Akbulut, Amanda McCroskery +17

Recent generative AI systems have demonstrated more advanced persuasive capabilities and are increasingly permeating areas of life where they can influence decision-making. Generat…

cs.CY2025

Evaluating Intra-firm LLM Alignment Strategies in Business Contexts

Noah Broestl, Benjamin Lange, Cristina Voinea +2

Instruction-tuned Large Language Models (LLMs) are increasingly deployed as AI Assistants in firms for support in cognitive tasks. These AI assistants carry embedded perspectives w…