1 citations · 1 across the 1 of their papers we have counts for
1 paper
Tadashi Kozuno, Yunhao Tang, Mark Rowland +5
Off-policy multi-step reinforcement learning algorithms consist of conservative and non-conservative algorithms: the former actively cut traces, whereas the latter do not. Recently…