How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies
arXiv:1512.02011
Abstract
Using deep neural nets as function approximator for reinforcement learning tasks have recently been shown to be very powerful for solving problems approaching real-world complexity. Using these results as a benchmark, we discuss the role that the discount factor may play in the quality of the learning process of a deep Q-network (DQN). When the discount factor progressively increases up to its final value, we empirically show that it is possible to significantly reduce the number of learning steps. When used in conjunction with a varying learning rate, we empirically show that it outperforms original DQN on several experiments. We relate this phenomenon with the instabilities of neural networks when they are used in an approximate Dynamic Programming setting. We also describe the possibility to fall within a local optimum during the learning process, thus connecting our discussion with the exploration/exploitation dilemma.
NIPS 2015 Deep Reinforcement Learning Workshop
References in corpus (2)
Cited by in corpus (11)
- Multi-Agent Deep Reinforcement Learning for Dynamic Power Allocation in Wireless Networks
- Comparison of Deep Reinforcement Learning and Model Predictive Control for Adaptive Cruise Control
- Hyperbolic Discounting and Learning over Multiple Horizons
- Personalized HeartSteps: A Reinforcement Learning Algorithm for Optimizing Physical Activity
- Beyond the One Step Greedy Approach in Reinforcement Learning
- Fast Efficient Hyperparameter Tuning for Policy Gradients
- Sample-Efficient Automated Deep Reinforcement Learning
- Neural Network Based Model Predictive Control for an Autonomous Vehicle
- Beyond Exponentially Discounted Sum: Automatic Learning of Return Function
- Online Meta-learning by Parallel Algorithm Competition
- MDP Playground: An Analysis and Debug Testbed for Reinforcement Learning