1 paper
Yujie Zhu, Charles A. Hepburn, Matthew Thorpe +1
In reinforcement learning with sparse rewards, demonstrations can accelerate learning, but determining when to imitate them remains challenging. We propose Smooth Policy Regularisa…