1 paper
Ziming Li, Sungjin Lee, Baolin Peng +5
Reinforcement Learning (RL) methods have emerged as a popular choice for training an efficient and effective dialogue policy. However, these methods suffer from sparse and unstable…