4 citations · 7 across the 14 of their papers we have counts for
1 paper · 1 filter
Lingwei Zhu, Takamitsu Matsubara
This paper aims to establish an entropy-regularized value-based reinforcement learning method that can ensure the monotonic improvement of policies at each policy update. Unlike pr…