The Uncertainty Bellman Equation and Exploration
arXiv:1709.05380
Abstract
We consider the exploration/exploitation problem in reinforcement learning. For exploitation, it is well known that the Bellman equation connects the value at any time-step to the expected value at subsequent time-steps. In this paper we consider a similar \textit{uncertainty} Bellman equation (UBE), which connects the uncertainty at any time-step to the expected uncertainties at subsequent time-steps, thereby extending the potential exploratory benefit of a policy beyond individual time-steps. We prove that the unique fixed point of the UBE yields an upper bound on the variance of the posterior distribution of the Q-values induced by any policy. This bound can be much tighter than traditional count-based bonuses that compound standard deviation rather than variance. Importantly, and unlike several existing approaches to optimism, this method scales naturally to large systems with complex generalization. Substituting our UBE-exploration strategy for -greedy improves DQN performance on 51 out of 57 games in the Atari suite.
Cited by in corpus (20)
- A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges
- Exploration in Deep Reinforcement Learning: A Survey
- Exploration by Random Network Distillation
- Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
- RUDDER: Return Decomposition for Delayed Rewards
- Reinforcement Learning through Active Inference
- Autonomous Platoon Control with Integrated Deep Reinforcement Learning and Dynamic Programming
- Unicorn: Continual Learning with a Universal, Off-policy Agent
- Optimal Scheduling in IoT-Driven Smart Isolated Microgrids Based on Deep Reinforcement Learning
- Benchmarking Bonus-Based Exploration Methods on the Arcade Learning Environment
- Efficient exploration with Double Uncertain Value Networks
- Expert-Supervised Reinforcement Learning for Offline Policy Learning and Evaluation
- Assumed Density Filtering Q-learning
- Kalman meets Bellman: Improving Policy Evaluation through Value Tracking
- Pessimistic Value Iteration for Multi-Task Data Sharing in Offline Reinforcement Learning
- Uncertainty-sensitive Learning and Planning with Ensembles
- Deep Model-Based Reinforcement Learning via Estimated Uncertainty and Conservative Policy Optimization
- Optimistic Proximal Policy Optimization
- Nonparametric Additive Value Functions: Interpretable Reinforcement Learning with an Application to Surgical Recovery
- Parameterized Indexed Value Function for Efficient Exploration in Reinforcement Learning