2 citations · 4 across the 2 of their papers we have counts for
2 papers
stat.ML2017★ 2 cited
Unifying Value Iteration, Advantage Learning, and Dynamic Policy Programming
Tadashi Kozuno, Eiji Uchibe, Kenji Doya
Approximate dynamic programming algorithms, such as approximate value iteration, have been successfully applied to many complex reinforcement learning tasks, and a better approxima…
cs.LG2017★ 2 cited
Online Meta-learning by Parallel Algorithm Competition
Stefan Elfwing, Eiji Uchibe, Kenji Doya
The efficiency of reinforcement learning algorithms depends critically on a few meta-parameters that modulates the learning updates and the trade-off between exploration and exploi…