Weighted Entropy Modification for Soft Actor-Critic
arXiv:2011.09083
Abstract
We generalize the existing principle of the maximum Shannon entropy in reinforcement learning (RL) to weighted entropy by characterizing the state-action pairs with some qualitative weights, which can be connected with prior knowledge, experience replay, and evolution process of the policy. We propose an algorithm motivated for self-balancing exploration with the introduced weight function, which leads to state-of-the-art performance on Mujoco tasks despite its simplicity in implementation.
References in corpus (3)
- Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
- Tsallis Reinforcement Learning: A Unified Framework for Maximum Entropy Reinforcement Learning
- Off-Policy Actor-Critic in an Ensemble: Achieving Maximum General Entropy and Effective Environment Exploration in Deep Reinforcement Learning