10 citations · 12 across the 13 of their papers we have counts for
5 papers · 1 filter
A dataset of rated conceptual arguments
Emery Cooper, Caspar Oesterheld, Linh Chi Nguyen +2
Large language models have improved rapidly on tasks with verifiable answers, such as mathematics and programming. Much less is known about their ability to reason about what we ca…
Implementing surrogate goals for safer bargaining in LLM-based agents
Caspar Oesterheld, Maxime Riché, Filip Sondej +2
Surrogate goals have been proposed as a strategy for reducing risks from bargaining failures. A surrogate goal is goal that a principal can give an AI agent and that deflects any t…
Observation Interference in Partially Observable Assistance Games
Scott Emmons, Caspar Oesterheld, Vincent Conitzer +1
We study partially observable assistance games (POAGs), a model of the human-AI value alignment problem which allows the human and the AI assistant to have partial observations. Mo…
Can CDT rationalise the ex ante optimal policy via modified anthropics?
Emery Cooper, Caspar Oesterheld, Vincent Conitzer
In Newcomb's problem, causal decision theory (CDT) recommends two-boxing and thus comes apart from evidential decision theory (EDT) and ex ante policy optimisation (which prescribe…
Recursive Joint Simulation in Games
Vojtech Kovarik, Caspar Oesterheld, Vincent Conitzer
Game-theoretic dynamics between AI agents could differ from traditional human-human interactions in various ways. One such difference is that it may be possible to accurately simul…