5 citations · 5 across the 1 of their papers we have counts for
1 paper · 1 filter
Domenic Donato, Lei Yu, Wang Ling +1
We introduce a new distributed policy gradient algorithm and show that it outperforms existing reward-aware training procedures such as REINFORCE, minimum risk training (MRT) and p…