Publications (19)
The Ethics of Advanced AI Assistants
Iason Gabriel, Arianna Manzini, Geoff Keeling +54
This paper focuses on the opportunities and the ethical and societal risks posed by advanced AI assistants. We define advanced AI assistants as artificial agents with natural langu…
Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points
Roberta Rocca, Sami Boukortt, Geoff Keeling +1
The paper proposes a new two‑player dialogue game, the Epistemic Asymmetry Schelling Task (EAST), to assess Theory of Mind and epistemic tracking abilities in large language models…
Chuck, Wilson and the emergence of artificial minds in human-AI conversations
Geoff Keeling, Winnie Street
Large Language Models (LLMs) can simulate person-like things which at least appear to have stable behavioural and psychological dispositions. Call these things characters. Are char…
LLMs achieve adult human performance on higher-order theory of mind tasks
Winnie Street, John Oliver Siy, Geoff Keeling +7
This paper examines the extent to which large language models (LLMs) have developed higher-order theory of mind (ToM); the human ability to reason about multiple mental and emotion…
Should agentic conversational AI change how we think about ethics? Characterising an interactional ethics centred on respect
Lize Alberts, Geoff Keeling, Amanda McCroskery
With the growing popularity of conversational agents based on large language models (LLMs), we need to ensure their behaviour is ethical and appropriate. Work in this area largely…
Inducing language models to assert their own consciousness restores human beliefs and values
Junsol Kim, Winnie Street, Roberta Rocca +4
The paper investigates how safety fine‑tuning of large language models reduces their tendency to attribute consciousness to themselves, animals, and objects, and shows that reversi…
Epistemic Trust as a Mechanism for Ethics Integration: Failure Modes and Design Principles from 70 Moral Imagination Workshops
Benjamin Lange, Geoff Keeling, Kyle Pedersen +4
Bottom-up responsible innovation initiatives seek to empower technology development teams to engage in ethical reflection, yet such interventions frequently fail to achieve practit…
We Need Accountability in Human-AI Agent Relationships
Benjamin Lange, Geoff Keeling, Arianna Manzini +1
We argue that accountability mechanisms are needed in human-AI agent relationships to ensure alignment with user and societal interests. We propose a framework according to which A…
Engaging Engineering Teams Through Moral Imagination: A Bottom-Up Approach for Responsible Innovation and Ethical Culture Change in Technology Companies
Benjamin Lange, Geoff Keeling, Amanda McCroskery +5
We propose a "Moral Imagination" methodology to facilitate a culture of responsible innovation for engineering and product teams in technology companies. Our approach has been oper…
On the attribution of confidence to large language models
Geoff Keeling, Winnie Street
Credences are mental states corresponding to degrees of confidence in propositions. Attribution of credences to Large Language Models (LLMs) is commonplace in the empirical literat…
We Need a New Ethics for a World of AI Agents
Iason Gabriel, Geoff Keeling, Arianna Manzini +1
The deployment of capable AI agents raises fresh questions about safety, human-machine relationships and social coordination. We argue for greater engagement by scientists, scholar…
Architecting Trust in Artificial Epistemic Agents
Nahema Marchal, Stephanie Chan, Matija Franklin +5
Large language models increasingly function as epistemic agents -- entities that can 1) autonomously pursue epistemic goals and 2) actively shape our shared knowledge environment.…
Deflating Deflationism: A Critical Perspective on Debunking Arguments Against LLM Mentality
Alex Grzankowski, Geoff Keeling, Henry Shevlin +1
Many people feel compelled to interpret, describe, and respond to Large Language Models (LLMs) as if they possess inner mental lives similar to our own. Responses to this phenomeno…
Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs
Junsol Kim, Winnie Street, Roberta Rocca +4
Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to e…
On the Opportunities and Risks of Foundation Models
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli +111
AI is undergoing a paradigm shift with the rise of models (e.g., BERT, DALL-E, GPT-3) that are trained on broad data at scale and are adaptable to a wide range of downstream tasks.…
Algorithmic Bias, Generalist Models,and Clinical Medicine
Geoff Keeling
The technical landscape of clinical machine learning is shifting in ways that destabilize pervasive assumptions about the nature and causes of algorithmic bias. On one hand, the do…
Can LLMs make trade-offs involving stipulated pain and pleasure states?
Geoff Keeling, Winnie Street, Martyna Stachaczyk +7
Pleasure and pain play an important role in human decision making by providing a common currency for resolving motivational conflicts. While Large Language Models (LLMs) can genera…
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
Seliem El-Sayed, Canfer Akbulut, Amanda McCroskery +17
Recent generative AI systems have demonstrated more advanced persuasive capabilities and are increasingly permeating areas of life where they can influence decision-making. Generat…
Evaluating Intra-firm LLM Alignment Strategies in Business Contexts
Noah Broestl, Benjamin Lange, Cristina Voinea +2
Instruction-tuned Large Language Models (LLMs) are increasingly deployed as AI Assistants in firms for support in cognitive tasks. These AI assistants carry embedded perspectives w…