activity
20172022
most citedBehavior Regularized Offline Reinforcement Learning

248 citations · 1.1k across the 25 of their papers we have counts for

collaborators

37 papers

cs.LG20227 cited

Why So Pessimistic? Estimating Uncertainties for Offline RL through Ensembles, and Why Their Independence Matters

Seyed Kamyar Seyed Ghasemipour, Shixiang Shane Gu, Ofir Nachum

Motivated by the success of ensembles for uncertainty estimation in supervised learning, we take a renewed look at how ensembles of -functions can be leveraged as the primary so…

cs.LG20224 cited

Chain of Thought Imitation with Procedure Cloning

Mengjiao Yang, Dale Schuurmans, Pieter Abbeel +1

Imitation learning aims to extract high-performance policies from logged demonstrations of expert behavior. It is common to frame imitation learning as a supervised learning proble…

cs.LG20215 cited

TRAIL: Near-Optimal Imitation Learning with Suboptimal Data

Mengjiao Yang, Sergey Levine, Ofir Nachum

The aim in imitation learning is to learn effective policies by utilizing near-optimal expert demonstrations. However, high-quality demonstrations from human experts can be expensi…

cs.LG20211 cited

Policy Gradients Incorporating the Future

David Venuto, Elaine Lau, Doina Precup +1

Reasoning about the future -- understanding how decisions in the present time affect outcomes in the future -- is one of the central challenges for reinforcement learning (RL), esp…

cs.LG2021

Provable Representation Learning for Imitation with Contrastive Fourier Features

Ofir Nachum, Mengjiao Yang

In imitation learning, it is common to learn a behavior policy to match an unknown target policy via max-likelihood training on a collected set of target demonstrations. In this wo…

cs.LG2021

Autoregressive Dynamics Models for Offline Policy Evaluation and Optimization

Michael R. Zhang, Tom Le Paine, Ofir Nachum +4

Standard dynamics models for continuous control make use of feedforward computation to predict the conditional distribution of next state and reward given current state and action…