22 citations · 53 across the 7 of their papers we have counts for
13 papers
Understanding Self-Predictive Learning for Reinforcement Learning
Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond +13
We study the learning dynamics of self-predictive learning for reinforcement learning, a family of algorithms that learn representations by minimizing the prediction error of their…
Universal Off-Policy Evaluation
Yash Chandak, Scott Niekum, Bruno Castro da Silva +3
When faced with sequential decision-making problems, it is often useful to be able to predict what would happen if decisions were made using a new policy. Those predictions must of…
High-Confidence Off-Policy (or Counterfactual) Variance Estimation
Yash Chandak, Shiv Shankar, Philip S. Thomas
Many sequential decision-making systems leverage data collected using prior policies to propose a new policy. For critical applications, it is important that high-confidence guaran…
Towards Safe Policy Improvement for Non-Stationary MDPs
Yash Chandak, Scott M. Jordan, Georgios Theocharous +2
Many real-world sequential decision-making problems involve critical systems with financial risks and human-life risks. While several works in the past have proposed methods that a…
Reinforcement Learning for Strategic Recommendations
Georgios Theocharous, Yash Chandak, Philip S. Thomas +1
Strategic recommendations (SR) refer to the problem where an intelligent agent observes the sequential behaviors and activities of users and decides when and how to interact with t…
Evaluating the Performance of Reinforcement Learning Algorithms
Scott M. Jordan, Yash Chandak, Daniel Cohen +2
Performance evaluations are critical for quantifying algorithmic advances in reinforcement learning. Recent reproducibility analyses have shown that reported performance results ar…