2 citations · 2 across the 1 of their papers we have counts for
1 paper
Zhengxu Hou, Bang Liu, Ruihui Zhao +4
For task-oriented dialog systems, training a Reinforcement Learning (RL) based Dialog Management module suffers from low sample efficiency and slow convergence speed due to the spa…