activity
20122022
most citedTemporal-Difference Networks

57 citations · 230 across the 22 of their papers we have counts for

collaborators
Showing cs.LGShow all

24 papers · 1 filter

cs.LG20223 cited

A History of Meta-gradient: Gradient Methods for Meta-learning

Richard S. Sutton

The history of meta-learning methods based on gradient descent is reviewed, focusing primarily on methods that adapt step-size (learning rate) meta-parameters.

cs.LG2021

Average-Reward Learning and Planning with Options

Yi Wan, Abhishek Naik, Richard S. Sutton

We extend the options framework for temporal abstraction in reinforcement learning from discounted Markov decision processes (MDPs) to average-reward MDPs. Our contributions includ…

cs.LG2021

An Empirical Comparison of Off-policy Prediction Learning Algorithms in the Four Rooms Environment

Sina Ghiassian, Richard S. Sutton

Many off-policy prediction learning algorithms have been proposed in the past decade, but it remains unclear which algorithms learn faster than others. We empirically compare 11 of…

cs.LG2021

An Empirical Comparison of Off-policy Prediction Learning Algorithms on the Collision Task

Sina Ghiassian, Richard S. Sutton

Off-policy prediction -- learning the value function for one policy from data generated while following another policy -- is one of the most challenging subproblems in reinforcemen…

cs.LG2021

Scalable Online Recurrent Learning Using Columnar Neural Networks

Khurram Javed, Martha White, Rich Sutton

Structural credit assignment for recurrent learning is challenging. An algorithm called RTRL can compute gradients for recurrent networks online but is computationally intractable…

cs.LG2021

Does the Adam Optimizer Exacerbate Catastrophic Forgetting?

Dylan R. Ashley, Sina Ghiassian, Richard S. Sutton

Catastrophic forgetting remains a severe hindrance to the broad application of artificial neural networks (ANNs), however, it continues to be a poorly understood phenomenon. Despit…