activity
20162026
most citedExtrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations

30 citations · 111 across the 49 of their papers we have counts for

collaborators
Showing 2021Show all

9 papers · 1 filter

cs.LG2021

SOPE: Spectrum of Off-Policy Estimators

Christina J. Yuan, Yash Chandak, Stephen Giguere +2

Many sequential decision making problems are high-stakes and require off-policy evaluation (OPE) of a new policy using historical data collected using some other policy. One of the…

cs.LG2021★ 2 cited

You Only Evaluate Once: a Simple Baseline Algorithm for Offline RL

Wonjoon Goo, Scott Niekum

The goal of offline reinforcement learning (RL) is to find an optimal policy given prerecorded trajectories. Many current approaches customize existing off-policy RL algorithms, es…

cs.RO2021

Distributional Depth-Based Estimation of Object Articulation Models

Ajinkya Jain, Stephen Giguere, Rudolf Lioutikov +1

We propose a method that efficiently learns distributions over articulation model parameters directly from depth images without the need to know articulation model categories a pri…

cs.LG2021

On the Benefits of Inducing Local Lipschitzness for Robust Generative Adversarial Imitation Learning

Farzan Memarian, Abolfazl Hashemi, Scott Niekum +1

We explore methodologies to improve the robustness of generative adversarial imitation learning (GAIL) algorithms to observation noise. Towards this objective, we study the effect…

cs.AI2021★ 5 cited

Zero-shot Task Adaptation using Natural Language

Prasoon Goyal, Raymond J. Mooney, Scott Niekum

Imitation learning and instruction-following are two common approaches to communicate a user's intent to a learning agent. However, as the complexity of tasks grows, it could be be…

cs.LG2021

Adversarial Intrinsic Motivation for Reinforcement Learning

Ishan Durugkar, Mauricio Tec, Scott Niekum +1

Learning with an objective to minimize the mismatch with a reference distribution has been shown to be useful for generative modeling and imitation learning. In this paper, we inve…