activity
20122022
most citedTemporal-Difference Networks

57 citations · 230 across the 22 of their papers we have counts for

collaborators
Showing 2021Show all

6 papers · 1 filter

cs.LG2021

Average-Reward Learning and Planning with Options

Yi Wan, Abhishek Naik, Richard S. Sutton

We extend the options framework for temporal abstraction in reinforcement learning from discounted Markov decision processes (MDPs) to average-reward MDPs. Our contributions includ…

cs.LG2021

An Empirical Comparison of Off-policy Prediction Learning Algorithms in the Four Rooms Environment

Sina Ghiassian, Richard S. Sutton

Many off-policy prediction learning algorithms have been proposed in the past decade, but it remains unclear which algorithms learn faster than others. We empirically compare 11 of…

cs.LG2021

An Empirical Comparison of Off-policy Prediction Learning Algorithms on the Collision Task

Sina Ghiassian, Richard S. Sutton

Off-policy prediction -- learning the value function for one policy from data generated while following another policy -- is one of the most challenging subproblems in reinforcemen…

cs.AI2021

Planning with Expectation Models for Control

Katya Kudashkina, Yi Wan, Abhishek Naik +1

In model-based reinforcement learning (MBRL), Wan et al. (2019) showed conditions under which the environment model could produce the expectation of the next feature vector rather…

cs.LG2021

Scalable Online Recurrent Learning Using Columnar Neural Networks

Khurram Javed, Martha White, Rich Sutton

Structural credit assignment for recurrent learning is challenging. An algorithm called RTRL can compute gradients for recurrent networks online but is computationally intractable…

cs.LG2021

Does the Adam Optimizer Exacerbate Catastrophic Forgetting?

Dylan R. Ashley, Sina Ghiassian, Richard S. Sutton

Catastrophic forgetting remains a severe hindrance to the broad application of artificial neural networks (ANNs), however, it continues to be a poorly understood phenomenon. Despit…