Curiosity creates Diversity in Policy Search
arXiv:2212.03530 · doi:10.1145/3605782
Abstract
When searching for policies, reward-sparse environments often lack sufficient information about which behaviors to improve upon or avoid. In such environments, the policy search process is bound to blindly search for reward-yielding transitions and no early reward can bias this search in one direction or another. A way to overcome this is to use intrinsic motivation in order to explore new transitions until a reward is found. In this work, we use a recently proposed definition of intrinsic motivation, Curiosity, in an evolutionary policy search method. We propose Curiosity-ES, an evolutionary strategy adapted to use Curiosity as a fitness metric. We compare Curiosity with Novelty, a commonly used diversity metric, and find that Curiosity can generate higher diversity over full episodes without the need for an explicit diversity criterion and lead to multiple policies which find reward.
Transactions on Evolutionary Learning and Optimization. 2023
References in corpus (12)
- Robots that can adapt like animals
- DeepMind Control Suite
- First return, then explore
- Variational Intrinsic Control
- Dream to Control: Learning Behaviors by Latent Imagination
- Covariance Matrix Adaptation for the Rapid Illumination of Behavior Space
- A survey on intrinsic motivation in reinforcement learning
- Autonomous skill discovery with Quality-Diversity and Unsupervised Descriptors
- EvoJAX: Hardware-Accelerated Neuroevolution
- ARCH-Elites: Quality-Diversity for Urban Design
- Sparse Reward Exploration via Novelty Search and Emitters
- Combining Evolution and Deep Reinforcement Learning for Policy Search: a Survey