activity
20122022
most citedTemporal-Difference Networks

57 citations · 230 across the 22 of their papers we have counts for

collaborators
Showing cs.AIShow all

11 papers · 1 filter

cs.AI2021

Planning with Expectation Models for Control

Katya Kudashkina, Yi Wan, Abhishek Naik +1

In model-based reinforcement learning (MBRL), Wan et al. (2019) showed conditions under which the environment model could produce the expectation of the next feature vector rather…

cs.AI20204 cited

Document-editing Assistants and Model-based Reinforcement Learning as a Path to Conversational AI

Katya Kudashkina, Patrick M. Pilarski, Richard S. Sutton

Intelligent assistants that follow commands or answer simple questions, such as Siri and Google search, are among the most economically important applications of AI. Future convers…

cs.AI201927 cited

Discounted Reinforcement Learning Is Not an Optimization Problem

Abhishek Naik, Roshan Shariff, Niko Yasui +2

Discounted reinforcement learning is fundamentally incompatible with function approximation for control in continuing tasks. It is not an optimization problem in its usual formulat…

cs.AI20191 cited

Should All Temporal Difference Learning Use Emphasis?

Xiang Gu, Sina Ghiassian, Richard S. Sutton

Emphatic Temporal Difference (ETD) learning has recently been proposed as a convergent off-policy learning method. ETD was proposed mainly to address convergence issues of conventi…

cs.AI2018

Reactive Reinforcement Learning in Asynchronous Environments

Jaden B. Travnik, Kory W. Mathewson, Richard S. Sutton +1

The relationship between a reinforcement learning (RL) agent and an asynchronous environment is often ignored. Frequently used models of the interaction between an agent and its en…

cs.AI2018

Directly Estimating the Variance of the λ-Return Using Temporal-Difference Methods

Craig Sherstan, Brendan Bennett, Kenny Young +4

This paper investigates estimating the variance of a temporal-difference learning agent's update target. Most reinforcement learning methods use an estimate of the value function,…