90 citations · 194 across the 7 of their papers we have counts for
1 paper · 1 filter
Nino Vieillard, Tadashi Kozuno, Bruno Scherrer +3
Recent Reinforcement Learning (RL) algorithms making use of Kullback-Leibler (KL) regularization as a core component have shown outstanding performance. Yet, only little is underst…