activity
20162026
most citedData-Efficient Policy Evaluation Through Behavior Policy Search

9 citations · 41 across the 36 of their papers we have counts for

collaborators
Showing 2024Show all

7 papers · 1 filter

cs.RO2024

Reinforcement Learning Within the Classical Robotics Stack: A Case Study in Robot Soccer

Adam Labiosa, Zhihan Wang, Siddhant Agarwal +10

Robot decision-making in partially observable, real-time, dynamic, and multi-agent environments remains a difficult and unsolved challenge. Model-free reinforcement learning (RL) i…

cs.LG2024

Stable Offline Value Function Learning with Bisimulation-based Representations

Brahma S. Pavse, Yudong Chen, Qiaomin Xie +1

In reinforcement learning, offline value function learning is the procedure of using an offline dataset to estimate the expected discounted return from each state when taking actio…

cs.LG2024

Reinforcement Learning via Auxiliary Task Distillation

Abhinav Narayan Harish, Larry Heck, Josiah P. Hanna +2

We present Reinforcement Learning via Auxiliary Task Distillation (AuxDistill), a new method that enables reinforcement learning (RL) to perform long-horizon robot control problems…

cs.LG2024★ 1 cited

SaVeR: Optimal Data Collection Strategy for Safe Policy Evaluation in Tabular MDP

Subhojyoti Mukherjee, Josiah P. Hanna, Robert Nowak

In this paper, we study safe data collection for the purpose of policy evaluation in tabular Markov decision processes (MDPs). In policy evaluation, we are given a \textit{target}…

cs.LG2024

Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning

Subhojyoti Mukherjee, Josiah P. Hanna, Qiaomin Xie +1

We study learning to learn for the multi-task structured bandit problem where the goal is to learn a near-optimal algorithm that minimizes cumulative regret. The tasks share a comm…

cs.LG2024★ 1 cited

Adaptive Exploration for Data-Efficient General Value Function Evaluations

Arushi Jain, Josiah P. Hanna, Doina Precup

General Value Functions (GVFs) (Sutton et al., 2011) represent predictive knowledge in reinforcement learning. Each GVF computes the expected return for a given policy, based on a…