2 citations · 3 across the 4 of their papers we have counts for
4 papers
Reinforcement Learning via Auxiliary Task Distillation
Abhinav Narayan Harish, Larry Heck, Josiah P. Hanna +2
We present Reinforcement Learning via Auxiliary Task Distillation (AuxDistill), a new method that enables reinforcement learning (RL) to perform long-horizon robot control problems…
SaVeR: Optimal Data Collection Strategy for Safe Policy Evaluation in Tabular MDP
Subhojyoti Mukherjee, Josiah P. Hanna, Robert Nowak
In this paper, we study safe data collection for the purpose of policy evaluation in tabular Markov decision processes (MDPs). In policy evaluation, we are given a \textit{target}…
Multi-task Representation Learning for Pure Exploration in Bilinear Bandits
Subhojyoti Mukherjee, Qiaomin Xie, Josiah P. Hanna +1
We study multi-task representation learning for the problem of pure exploration in bilinear bandits. In bilinear bandits, an action takes the form of a pair of arms from two differ…
State-Action Similarity-Based Representations for Off-Policy Evaluation
Brahma S. Pavse, Josiah P. Hanna
In reinforcement learning, off-policy evaluation (OPE) is the problem of estimating the expected return of an evaluation policy given a fixed dataset that was collected by running…