4 papers · 1 filter
A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences
Jobst Heitzig, Ram Potham
This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage the power balance between h…
Unbiased Canonical Set-Valued Oracles Via Lattice Theory
Jobst Heitzig
An oracle that tells you the probability of some future event can change that very probability because you act on the answer. We argue that this performativity is OK as people cons…
Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
Jobst Heitzig, Ram Potham
Power is a key concept in AI safety: power-seeking as an instrumental goal, sudden or gradual disempowerment of humans, power balance in human-AI interaction and international AI g…
Non-maximizing policies that fulfill multi-criterion aspirations in expectation
Simon Dima, Simon Fischer, Jobst Heitzig +1
In dynamic programming and reinforcement learning, the policy for the sequential decision making of an agent in a stochastic environment is usually determined by expressing the goa…