2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Hong Liu, Zhijian Ou, Yi Huang +1
Recently, there has been progress in supervised funetuning pretrained GPT-2 to build end-to-end task-oriented dialog (TOD) systems. However, online reinforcement learning of a GPT-…