4 citations · 7 across the 14 of their papers we have counts for
3 papers · 1 filter
Geometric Value Iteration: Dynamic Error-Aware KL Regularization for Reinforcement Learning
Toshinori Kitamura, Lingwei Zhu, Takamitsu Matsubara
The recent boom in the literature on entropy-regularized reinforcement learning (RL) approaches reveals that Kullback-Leibler (KL) regularization brings advantages to RL algorithms…
Cautious Policy Programming: Exploiting KL Regularization in Monotonic Policy Improvement for Reinforcement Learning
Lingwei Zhu, Toshinori Kitamura, Takamitsu Matsubara
In this paper, we propose cautious policy programming (CPP), a novel value-based reinforcement learning (RL) algorithm that can ensure monotonic policy improvement during learning.…
Cautious Actor-Critic
Lingwei Zhu, Toshinori Kitamura, Takamitsu Matsubara
The oscillating performance of off-policy learning and persisting errors in the actor-critic (AC) setting call for algorithms that can conservatively learn to suit the stability-cr…