papers

Publications (9)

cs.AI2025

Automated Reward Design for Gran Turismo

Michel Ma, Takuma Seno, Kaushik Subramanian +3

When designing reinforcement learning (RL) agents, a designer communicates the desired agent behavior through the definition of reward functions - numerical feedback given to the a…

cs.AI2016

Introspective Agents: Confidence Measures for General Value Functions

Craig Sherstan, Adam White, Marlos C. Machado +1

Agents of general intelligence deployed in real-world scenarios must adapt to ever-changing environmental conditions. While such adaptive agents may leverage engineered knowledge,…

cs.LG2022

Value Function Decomposition for Iterative Design of Reinforcement Learning Agents

James MacGlashan, Evan Archer, Alisa Devlic +4

Designing reinforcement learning (RL) agents is typically a difficult process that requires numerous design iterations. Learning can fail for a multitude of reasons, and standard R…

cs.AI2017

Communicative Capital for Prosthetic Agents

Patrick M. Pilarski, Richard S. Sutton, Kory W. Mathewson +3

This work presents an overarching perspective on the role that machine intelligence can play in enhancing human abilities, especially those that have been diminished due to injury…

cs.LG2020

Gamma-Nets: Generalizing Value Estimation over Timescale

Craig Sherstan, Shibhansh Dohare, James MacGlashan +2

We present -nets, a method for generalizing value function estimation over timescale. By using the timescale as one of the estimator's inputs we can estimate value for arbitrar…

cs.LG2020

Work in Progress: Temporally Extended Auxiliary Tasks

Craig Sherstan, Bilal Kartal, Pablo Hernandez-Leal +1

Predictive auxiliary tasks have been shown to improve performance in numerous reinforcement learning works, however, this effect is still not well understood. The primary purpose o…

cs.AI2026

Coachable agents for interactive gameplay

Roberto Capobianco, Harm van Seijen, Nolan D. Bard +39

Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation m…

cs.AI2018

Directly Estimating the Variance of the λ-Return Using Temporal-Difference Methods

Craig Sherstan, Brendan Bennett, Kenny Young +4

This paper investigates estimating the variance of a temporal-difference learning agent's update target. Most reinforcement learning methods use an estimate of the value function,…

cs.LG2018

Accelerating Learning in Constructive Predictive Frameworks with the Successor Representation

Craig Sherstan, Marlos C. Machado, Patrick M. Pilarski

Here we propose using the successor representation (SR) to accelerate learning in a constructive knowledge system based on general value functions (GVFs). In real-world settings li…