Learning values across many orders of magnitude
arXiv:1602.07714
Abstract
Most learning algorithms are not invariant to the scale of the function that is being approximated. We propose to adaptively normalize the targets used in learning. This is useful in value-based reinforcement learning, where the magnitude of appropriate value approximations can change over time when we update the policy of behavior. Our main motivation is prior work on learning to play Atari games, where the rewards were all clipped to a predetermined range. This clipping facilitates learning across many different games with a single learning algorithm, but a clipped reward function can result in qualitatively different behavior. Using the adaptive normalization we can remove this domain-specific heuristic without diminishing overall performance.
Paper accepted for publication at NIPS 2016. This version includes the appendix
References in corpus (5)
Cited by in corpus (29)
- Density estimation using Real NVP
- Deep Reinforcement Learning: An Overview
- Comparison of Deep Reinforcement Learning and Model Predictive Control for Adaptive Cruise Control
- Hybrid Reward Architecture for Reinforcement Learning
- A Survey of Deep Reinforcement Learning in Video Games
- Trial without Error: Towards Safe Reinforcement Learning via Human Intervention
- Review: Deep Learning in Electron Microscopy
- Towards V2I Age-aware Fairness Access: A DQN Based Intelligent Vehicular Node Training and Test Method
- Observe and Look Further: Achieving Consistent Performance on Atari
- Comparative analysis of machine learning methods for active flow control
- Longitudinal Dynamic versus Kinematic Models for Car-Following Control Using Deep Reinforcement Learning
- Networked Multiagent Safe Reinforcement Learning for Low-carbon Demand Management in Distribution Network
- Re-evaluating Evaluation
- Multi-task Deep Reinforcement Learning with PopArt
- Sample-Efficient Imitation Learning via Generative Adversarial Nets
- Lipschitzness Is All You Need To Tame Off-policy Generative Adversarial Imitation Learning
- Policy Gradient With Value Function Approximation For Collective Multiagent Planning
- Multi-Advisor Reinforcement Learning
- Muesli: Combining Improvements in Policy Optimization
- Analysis of Agent Expertise in Ms. Pac-Man using Value-of-Information-based Policies
- Computational Performance of Deep Reinforcement Learning to find Nash Equilibria
- Model-Free Reinforcement Learning for Financial Portfolios: A Brief Survey
- Separation of Concerns in Reinforcement Learning
- Towards continuous control of flippers for a multi-terrain robot using deep reinforcement learning
- Automata-Guided Hierarchical Reinforcement Learning for Skill Composition
- ANS: Adaptive Network Scaling for Deep Rectifier Reinforcement Learning Models
- Deep Q-Network for Angry Birds
- A Closer Look at Advantage-Filtered Behavioral Cloning in High-Noise Datasets
- Efficient Reinforcement Learning in Resource Allocation Problems Through Permutation Invariant Multi-task Learning