1 paper
Xuecheng Niu, Akinori Ito, Takashi Nose
Training task-oriented dialog agents based on reinforcement learning is time-consuming and requires a large number of interactions with real users. How to grasp dialog policy withi…