1 paper
Yangyang Zhao, Zhenyu Wang, Mehdi Dastani +1
Training a dialogue policy using deep reinforcement learning requires a lot of exploration of the environment. The amount of wasted invalid exploration makes their learning ineffic…