30 citations · 111 across the 49 of their papers we have counts for
9 papers · 1 filter
SOPE: Spectrum of Off-Policy Estimators
Christina J. Yuan, Yash Chandak, Stephen Giguere +2
Many sequential decision making problems are high-stakes and require off-policy evaluation (OPE) of a new policy using historical data collected using some other policy. One of the…
You Only Evaluate Once: a Simple Baseline Algorithm for Offline RL
Wonjoon Goo, Scott Niekum
The goal of offline reinforcement learning (RL) is to find an optimal policy given prerecorded trajectories. Many current approaches customize existing off-policy RL algorithms, es…
Distributional Depth-Based Estimation of Object Articulation Models
Ajinkya Jain, Stephen Giguere, Rudolf Lioutikov +1
We propose a method that efficiently learns distributions over articulation model parameters directly from depth images without the need to know articulation model categories a pri…
On the Benefits of Inducing Local Lipschitzness for Robust Generative Adversarial Imitation Learning
Farzan Memarian, Abolfazl Hashemi, Scott Niekum +1
We explore methodologies to improve the robustness of generative adversarial imitation learning (GAIL) algorithms to observation noise. Towards this objective, we study the effect…
Zero-shot Task Adaptation using Natural Language
Prasoon Goyal, Raymond J. Mooney, Scott Niekum
Imitation learning and instruction-following are two common approaches to communicate a user's intent to a learning agent. However, as the complexity of tasks grows, it could be be…
Adversarial Intrinsic Motivation for Reinforcement Learning
Ishan Durugkar, Mauricio Tec, Scott Niekum +1
Learning with an objective to minimize the mismatch with a reference distribution has been shown to be useful for generative modeling and imitation learning. In this paper, we inve…