12 papers
A dataset of rated conceptual arguments
Emery Cooper, Caspar Oesterheld, Linh Chi Nguyen +2
The paper introduces a dataset of 951 expert‑rated argumentative critiques on 442 position texts covering AI safety, decision theory, ethics, and politics, and uses it to benchmark…
Recursive Joint Simulation in Games
Vojtech Kovarik, Caspar Oesterheld, Vincent Conitzer
Game-theoretic dynamics between AI agents could differ from traditional human-human interactions in various ways. One such difference is that it may be possible to accurately simul…
Implementing surrogate goals for safer bargaining in LLM-based agents
Caspar Oesterheld, Maxime Riché, Filip Sondej +2
Surrogate goals have been proposed as a strategy for reducing risks from bargaining failures. A surrogate goal is goal that a principal can give an AI agent and that deflects any t…
How Many Votes is a Lie Worth? Measuring Strategyproofness through Resource Augmentation
Ratip Emin Berker, Vincent Conitzer, Eden Hartman +2
It is well known, by the Gibbard-Satterthwaite Theorem, that when there are more than two candidates, any non-dictatorial voting rule can be manipulated by untruthful voters. But h…
Choosing What Game to Play without Selecting Equilibria: Inferring Safe (Pareto) Improvements in Binary Constraint Structures
Caspar Oesterheld, Vincent Conitzer
We consider a setting in which a principal gets to choose which game from some given set is played by a group of agents. The principal would like to choose a game that favors one o…
Promises Made, Promises Kept: Safe Pareto Improvements via Ex Post Verifiable Commitments
Nathaniel Sauerberg, Caspar Oesterheld
A safe Pareto improvement (SPI) [41] is a modification of a game that leaves all players better off with certainty. SPIs are typically proven under qualitative assumptions about th…