activity
20172022
most citedConservative Q-Learning for Offline Reinforcement Learning

538 citations · 1k across the 11 of their papers we have counts for

collaborators

29 papers

cs.LG20221 cited

Oracle Inequalities for Model Selection in Offline Reinforcement Learning

Jonathan N. Lee, George Tucker, Ofir Nachum +2

In offline reinforcement learning (RL), a learner leverages prior logged data to learn a good policy without interacting with the environment. A major challenge in applying such me…

cs.LG2021

Autoregressive Dynamics Models for Offline Policy Evaluation and Optimization

Michael R. Zhang, Tom Le Paine, Ofir Nachum +4

Standard dynamics models for continuous control make use of feedforward computation to predict the conditional distribution of next state and reward given current state and action…

cs.LG202123 cited

Benchmarks for Deep Off-Policy Evaluation

Justin Fu, Mohammad Norouzi, Ofir Nachum +10

Off-policy evaluation (OPE) holds the promise of being able to leverage large, offline datasets for both evaluating and selecting complex policies for decision making. The ability…

cs.LG20208 cited

Offline Policy Selection under Uncertainty

Mengjiao Yang, Bo Dai, Ofir Nachum +2

The presence of uncertainty in policy evaluation significantly complicates the process of policy ranking and selection in real-world settings. We formally consider offline policy s…

cs.LG2020

RL Unplugged: A Suite of Benchmarks for Offline Reinforcement Learning

Caglar Gulcehre, Ziyu Wang, Alexander Novikov +15

Offline methods for reinforcement learning have a potential to help bridge the gap between reinforcement learning research and real-world applications. They make it possible to lea…

cs.LG2020

DisARM: An Antithetic Gradient Estimator for Binary Latent Variables

Zhe Dong, Andriy Mnih, George Tucker

Training models with discrete latent variables is challenging due to the difficulty of estimating the gradients accurately. Much of the recent progress has been achieved by taking…