Bad Universal Priors and Notions of Optimality
arXiv:1510.04931
Abstract
A big open question of algorithmic information theory is the choice of the universal Turing machine (UTM). For Kolmogorov complexity and Solomonoff induction we have invariance theorems: the choice of the UTM changes bounds only by a constant. For the universally intelligent agent AIXI (Hutter, 2005) no invariance theorem is known. Our results are entirely negative: we discuss cases in which unlucky or adversarial choices of the UTM cause AIXI to misbehave drastically. We show that Legg-Hutter intelligence and thus balanced Pareto optimality is entirely subjective, and that every policy is Pareto optimal in the class of all computable environments. This undermines all existing optimality properties for AIXI. While it may still serve as a gold standard for AI, our results imply that AIXI is a relative theory, dependent on the choice of the UTM.
COLT 2015
Cited by in corpus (9)
- Scalable agent alignment via reward modeling: a research direction
- Thompson Sampling is Asymptotically Optimal in General Environments
- Extending Environments To Measure Self-Reflection In Reinforcement Learning
- On the Computability of Solomonoff Induction and Knowledge-Seeking
- Reward-Punishment Symmetric Universal Intelligence
- A Formal Solution to the Grain of Truth Problem
- On the Computability of AIXI
- Self-Modification of Policy and Utility Function in Rational Agents
- Curiosity Killed or Incapacitated the Cat and the Asymptotically Optimal Agent