Muesli: Combining Improvements in Policy Optimization
arXiv:2104.06159
Abstract
We propose a novel policy update that combines regularized policy optimization with model learning as an auxiliary loss. The update (henceforth Muesli) matches MuZero's state-of-the-art performance on Atari. Notably, Muesli does so without using deep search: it acts directly with a policy network and has computation speed comparable to model-free baselines. The Atari results are complemented by extensive ablations, and by additional results on continuous control and 9x9 Go.
References in corpus (35)
- Adam: A Method for Stochastic Optimization
- Decoupled Weight Decay Regularization
- Trust Region Policy Optimization
- Stochastic Backpropagation and Approximate Inference in Deep Generative Models
- Soft Actor-Critic Algorithms and Applications
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
- Recurrent World Models Facilitate Policy Evolution
- Rainbow: Combining Improvements in Deep Reinforcement Learning
- Normalizing Flows for Probabilistic Modeling and Inference
- Reinforcement Learning with Unsupervised Auxiliary Tasks
- Implicit Quantile Networks for Distributional Reinforcement Learning
- Maximum a Posteriori Policy Optimisation
- Sample Efficient Actor-Critic with Experience Replay
- Imagination-Augmented Agents for Deep Reinforcement Learning
- Meta-Gradient Reinforcement Learning
- On the Theory of Policy Gradient Methods: Optimality, Approximation, and Distribution Shift
- When to use parametric models in reinforcement learning?
- Learning values across many orders of magnitude
- Observe and Look Further: Achieving Consistent Performance on Atari
- Learning model-based planning from scratch
- Phasic Policy Gradient
- Bootstrap Latent-Predictive Representations for Multitask Reinforcement Learning
- Online and Offline Reinforcement Learning by Planning with a Learned Model
- TreeQN and ATreeC: Differentiable Tree-Structured Models for Deep Reinforcement Learning
- Mirror Descent Policy Optimization
- On Inductive Biases in Deep Reinforcement Learning
- Causally Correct Partial Models for Reinforcement Learning
- Combining Q-Learning and Search with Amortized Value Estimates
- On the role of planning in model-based deep reinforcement learning
- Learning to Predict Independent of Span
- Learning and Planning in Complex Action Spaces
- General non-linear Bellman equations
- Podracer architectures for scalable Reinforcement Learning
- Local Search for Policy Iteration in Continuous Control