1 paper
Yangyang Zhao, Ben Niu, Libo Qin +1
Deep Reinforcement Learning (DRL) is widely used in task-oriented dialogue systems to optimize dialogue policy, but it struggles to balance exploration and exploitation due to the…