3 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.LG2021★ 3 cited
Preferential Temporal Difference Learning
Nishanth Anand, Doina Precup
Temporal-Difference (TD) learning is a general and very useful tool for estimating the value function of a given policy, which in turn is required to find good policies. Generally…
cs.LG2019★ 2 cited
Recurrent Value Functions
Pierre Thodoroff, Nishanth Anand, Lucas Caccia +2
Despite recent successes in Reinforcement Learning, value-based methods often suffer from high variance hindering performance. In this paper, we illustrate this in a continuous con…