Taming the Noise in Reinforcement Learning via Soft Updates
arXiv:1512.08562
Abstract
Model-free reinforcement learning algorithms, such as Q-learning, perform poorly in the early stages of learning in noisy environments, because much effort is spent unlearning biased estimates of the state-action value function. The bias results from selecting, among several noisy estimates, the apparent optimum, which may actually be suboptimal. We propose G-learning, a new off-policy learning algorithm that regularizes the value estimates by penalizing deterministic policies in the beginning of the learning process. We show that this method reduces the bias of the value-function estimation, leading to faster convergence to the optimal value and the optimal policy. Moreover, G-learning enables the natural incorporation of prior domain knowledge, when available. The stochastic nature of G-learning also makes it avoid some exploration costs, a property usually attributed only to on-policy algorithms. We illustrate these ideas in several examples, where G-learning results in significant improvements of the convergence rate and the cost of the learning process.
References in corpus (1)
Cited by in corpus (28)
- Reinforcement Learning with Deep Energy-Based Policies
- Equivalence Between Policy Gradients and Soft Q-Learning
- If MaxEnt RL is the Answer, What is the Question?
- Maximum Entropy RL (Provably) Solves Some Robust RL Problems
- Deep Reinforcement Learning amidst Lifelong Non-Stationarity
- Offline Reinforcement Learning with Fisher Divergence Critic Regularization
- Sequence Tutor: Conservative Fine-Tuning of Sequence Generation Models with KL-control
- Action and Perception as Divergence Minimization
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- On the model-based stochastic value gradient for continuous reinforcement learning
- Self-Imitation Learning via Generalized Lower Bound Q-learning
- Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient Exploration
- A Regularized Opponent Model with Maximum Entropy Objective
- Representing preorders with injective monotones
- Hindsight Expectation Maximization for Goal-conditioned Reinforcement Learning
- Information Theoretic Model Predictive Q-Learning
- Soft Policy Gradient Method for Maximum Entropy Deep Reinforcement Learning
- A short variational proof of equivalence between policy gradients and soft Q learning
- Finding the Near Optimal Policy via Adaptive Reduced Regularization in MDPs
- A Relation Analysis of Markov Decision Process Frameworks
- Recurrent Value Functions
- Market Self-Learning of Signals, Impact and Optimal Trading: Invisible Hand Inference with Free Energy
- Market-making with reinforcement-learning (SAC)
- SoftDICE for Imitation Learning: Rethinking Off-policy Distribution Matching
- Tomography Based Learning for Load Distribution through Opaque Networks
- Adaptive Symmetric Reward Noising for Reinforcement Learning
- Can Q-learning solve Multi Armed Bantids?
- Parameterized MDPs and Reinforcement Learning Problems -- A Maximum Entropy Principle Based Framework