activity
20162026
most citedData-Efficient Policy Evaluation Through Behavior Policy Search

9 citations · 41 across the 36 of their papers we have counts for

collaborators
Showing 2022Show all

5 papers · 1 filter

cs.LG2022

Safe Evaluation For Offline Learning: Are We Ready To Deploy?

Hager Radi, Josiah P. Hanna, Peter Stone +1

The world currently offers an abundance of data in multiple domains, from which we can learn reinforcement learning (RL) policies without further interaction with the environment.…

cs.LG2022★ 1 cited

Scaling Marginalized Importance Sampling to High-Dimensional State-Spaces via State Abstraction

Brahma S. Pavse, Josiah P. Hanna

We consider the problem of off-policy evaluation (OPE) in reinforcement learning (RL), where the goal is to estimate the performance of an evaluation policy, , using a fixed d…

cs.LG2022

A Joint Imitation-Reinforcement Learning Framework for Reduced Baseline Regret

Sheelabhadra Dey, Sumedh Pendurkar, Guni Sharon +1

In various control task domains, existing controllers provide a baseline level of performance that -- though possibly suboptimal -- should be maintained. Reinforcement learning (RL…

cs.LG2022★ 4 cited

Temporal Disentanglement of Representations for Improved Generalisation in Reinforcement Learning

Mhairi Dunion, Trevor McInroe, Kevin Sebastian Luck +2

Reinforcement Learning (RL) agents are often unable to generalise well to environment variations in the state space that were not observed during training. This issue is especially…

cs.LG2022★ 2 cited

ReVar: Strengthening Policy Evaluation via Reduced Variance Sampling

Subhojyoti Mukherjee, Josiah P. Hanna, Robert Nowak

This paper studies the problem of data collection for policy evaluation in Markov decision processes (MDPs). In policy evaluation, we are given a target policy and asked to estimat…