57 citations · 230 across the 22 of their papers we have counts for
4 papers · 1 filter
Communicative Capital for Prosthetic Agents
Patrick M. Pilarski, Richard S. Sutton, Kory W. Mathewson +3
This work presents an overarching perspective on the role that machine intelligence can play in enhancing human abilities, especially those that have been diminished due to injury…
A First Empirical Study of Emphatic Temporal Difference Learning
Sina Ghiassian, Banafsheh Rafiee, Richard S. Sutton
In this paper we present the first empirical study of the emphatic temporal-difference learning algorithm (ETD), comparing it with conventional temporal-difference learning, in par…
GQ() Quick Reference and Implementation Guide
Adam White, Richard S. Sutton
This document should serve as a quick reference for and guide to the implementation of linear GQ(), a gradient-based off-policy temporal-difference learning algorithm. Explanati…
Multi-step Off-policy Learning Without Importance Sampling Ratios
Ashique Rupam Mahmood, Huizhen Yu, Richard S. Sutton
To estimate the value functions of policies from exploratory data, most model-free off-policy algorithms rely on importance sampling, where the use of importance sampling ratios of…