9 citations · 30 across the 7 of their papers we have counts for
1 paper · 1 filter
Jorge A. Mendez, Alborz Geramifard, Mohammad Ghavamzadeh +1
Learning task-oriented dialog policies via reinforcement learning typically requires large amounts of interaction with users, which in practice renders such methods unusable for re…