ISL: A novel approach for deep exploration
arXiv:1909.06293
Abstract
In this article we explore an alternative approach to address deep exploration and we introduce the ISL algorithm, which is efficient at performing deep exploration. Similarly to maximum entropy RL, we derive the algorithm by augmenting the traditional RL objective with a novel regularization term. A distinctive feature of our approach is that, as opposed to other works that tackle the problem of deep exploration, in our derivation both the learning equations and the exploration-exploitation strategy are derived in tandem as the solution to a well-posed optimization problem whose minimization leads to the optimal value function. Empirically we show that our method exhibits state of the art performance on a range of challenging deep-exploration benchmarks.
References in corpus (8)
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Reinforcement Learning with Deep Energy-Based Policies
- Large-Scale Study of Curiosity-Driven Learning
- Count-Based Exploration with Neural Density Models
- Exploration by Random Network Distillation
- Bridging the Gap Between Value and Policy Based Reinforcement Learning
- SBEED: Convergent Reinforcement Learning with Nonlinear Function Approximation
- Behaviour Suite for Reinforcement Learning