activity
20162021
most citedPolicy Gradient Methods for Reinforcement Learning with Function Approximation and Action-Dependent Baselines

45 citations · 138 across the 16 of their papers we have counts for

collaborators
Showing cs.LGShow all

18 papers · 1 filter

cs.LG20215 cited

Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPs

Harsh Satija, Philip S. Thomas, Joelle Pineau +1

We study the problem of Safe Policy Improvement (SPI) under constraints in the offline Reinforcement Learning (RL) setting. We consider the scenario where: (i) we have a dataset co…

cs.LG2021

Universal Off-Policy Evaluation

Yash Chandak, Scott Niekum, Bruno Castro da Silva +3

When faced with sequential decision-making problems, it is often useful to be able to predict what would happen if decisions were made using a new policy. Those predictions must of…

cs.LG20211 cited

High-Confidence Off-Policy (or Counterfactual) Variance Estimation

Yash Chandak, Shiv Shankar, Philip S. Thomas

Many sequential decision-making systems leverage data collected using prior policies to propose a new policy. For critical applications, it is important that high-confidence guaran…

cs.LG2020

Towards Safe Policy Improvement for Non-Stationary MDPs

Yash Chandak, Scott M. Jordan, Georgios Theocharous +2

Many real-world sequential decision-making problems involve critical systems with financial risks and human-life risks. While several works in the past have proposed methods that a…

cs.LG20202 cited

Reinforcement Learning for Strategic Recommendations

Georgios Theocharous, Yash Chandak, Philip S. Thomas +1

Strategic recommendations (SR) refer to the problem where an intelligent agent observes the sequential behaviors and activities of users and decides when and how to interact with t…

cs.LG202019 cited

Evaluating the Performance of Reinforcement Learning Algorithms

Scott M. Jordan, Yash Chandak, Daniel Cohen +2

Performance evaluations are critical for quantifying algorithmic advances in reinforcement learning. Recent reproducibility analyses have shown that reported performance results ar…