45 citations · 138 across the 16 of their papers we have counts for
4 papers · 1 filter
Towards Practical Mean Bounds for Small Samples
My Phan, Philip S. Thomas, Erik Learned-Miller
Historically, to bound the mean for small sample sizes, practitioners have had to choose between using methods with unrealistic assumptions about the unknown distribution (e.g., Ga…
Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPs
Harsh Satija, Philip S. Thomas, Joelle Pineau +1
We study the problem of Safe Policy Improvement (SPI) under constraints in the offline Reinforcement Learning (RL) setting. We consider the scenario where: (i) we have a dataset co…
Universal Off-Policy Evaluation
Yash Chandak, Scott Niekum, Bruno Castro da Silva +3
When faced with sequential decision-making problems, it is often useful to be able to predict what would happen if decisions were made using a new policy. Those predictions must of…
High-Confidence Off-Policy (or Counterfactual) Variance Estimation
Yash Chandak, Shiv Shankar, Philip S. Thomas
Many sequential decision-making systems leverage data collected using prior policies to propose a new policy. For critical applications, it is important that high-confidence guaran…