From the 1 of 15 linked papers with an AI index.
5 papers · 1 filter
A dataset of rated conceptual arguments
Emery Cooper, Caspar Oesterheld, Linh Chi Nguyen +2
The paper introduces a dataset of 951 expert‑rated argumentative critiques on 442 position texts covering AI safety, decision theory, ethics, and politics, and uses it to benchmark…
Recursive Joint Simulation in Games
Vojtech Kovarik, Caspar Oesterheld, Vincent Conitzer
Game-theoretic dynamics between AI agents could differ from traditional human-human interactions in various ways. One such difference is that it may be possible to accurately simul…
Implementing surrogate goals for safer bargaining in LLM-based agents
Caspar Oesterheld, Maxime Riché, Filip Sondej +2
Surrogate goals have been proposed as a strategy for reducing risks from bargaining failures. A surrogate goal is goal that a principal can give an AI agent and that deflects any t…
Observation Interference in Partially Observable Assistance Games
Scott Emmons, Caspar Oesterheld, Vincent Conitzer +1
We study partially observable assistance games (POAGs), a model of the human-AI value alignment problem which allows the human and the AI assistant to have partial observations. Mo…
Can CDT rationalise the ex ante optimal policy via modified anthropics?
Emery Cooper, Caspar Oesterheld, Vincent Conitzer
In Newcomb's problem, causal decision theory (CDT) recommends two-boxing and thus comes apart from evidential decision theory (EDT) and ex ante policy optimisation (which prescribe…