activity
20242026
most citedMulti-Agent Risks from Advanced AI

10 citations · 12 across the 13 of their papers we have counts for

collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

A dataset of rated conceptual arguments

Emery Cooper, Caspar Oesterheld, Linh Chi Nguyen +2

Large language models have improved rapidly on tasks with verifiable answers, such as mathematics and programming. Much less is known about their ability to reason about what we ca…

cs.AI2026

Implementing surrogate goals for safer bargaining in LLM-based agents

Caspar Oesterheld, Maxime Riché, Filip Sondej +2

Surrogate goals have been proposed as a strategy for reducing risks from bargaining failures. A surrogate goal is goal that a principal can give an AI agent and that deflects any t…

cs.AI20241 cited

Observation Interference in Partially Observable Assistance Games

Scott Emmons, Caspar Oesterheld, Vincent Conitzer +1

We study partially observable assistance games (POAGs), a model of the human-AI value alignment problem which allows the human and the AI assistant to have partial observations. Mo…

cs.AI2024

Can CDT rationalise the ex ante optimal policy via modified anthropics?

Emery Cooper, Caspar Oesterheld, Vincent Conitzer

In Newcomb's problem, causal decision theory (CDT) recommends two-boxing and thus comes apart from evidential decision theory (EDT) and ex ante policy optimisation (which prescribe…

cs.AI2024

Recursive Joint Simulation in Games

Vojtech Kovarik, Caspar Oesterheld, Vincent Conitzer

Game-theoretic dynamics between AI agents could differ from traditional human-human interactions in various ways. One such difference is that it may be possible to accurately simul…