1 citations · 1 across the 1 of their papers we have counts for
1 paper
Hongru Wang, Huimin Wang, Zezhong Wang +1
Reinforcement Learning (RL) has been witnessed its potential for training a dialogue policy agent towards maximizing the accumulated rewards given from users. However, the reward c…