artificial intelligence

A Model-Free Universal AI

arXiv:2602.23242

summary

The paper proposes AIQI, a model-free reinforcement learning agent that uses universal induction over action-value functions and is proven to be asymptotically epsilon-optimal.

Abstract

In general reinforcement learning, all established optimal agents, including AIXI, are model-based, explicitly maintaining and using environment models. This paper introduces Universal AI with Q-Induction (AIQI), the first model-free agent proven to be asymptotically -optimal in general RL. AIQI performs universal induction over distributional action-value functions, instead of policies or environments like previous works. Under a grain of truth condition, we prove that AIQI is strong asymptotically -optimal and asymptotically -Bayes-optimal. We also apply our novel proof techniques to show asymptotic -optimality of Self-AIXI without any ad-hoc assumptions. Our results significantly expand the diversity of known universal agents.

42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026)

Topics & keywords

#reinforcement learning#model-free methods#universal AI#theoretical analysis#optimality guaranteesAIQIepsilon-optimaluniversal inductionaction-value functionSelf-AIXI
A Model-Free Universal AI · wovepaper