68 citations · 191 across the 9 of their papers we have counts for
1 paper · 1 filter
Jorge A. Mendez, Alborz Geramifard, Mohammad Ghavamzadeh +1
Learning task-oriented dialog policies via reinforcement learning typically requires large amounts of interaction with users, which in practice renders such methods unusable for re…