activity
20182021
most citedHyperparameter Selection for Offline Reinforcement Learning

30 citations · 57 across the 3 of their papers we have counts for

collaborators

10 papers

cs.LG2021

Autoregressive Dynamics Models for Offline Policy Evaluation and Optimization

Michael R. Zhang, Tom Le Paine, Ofir Nachum +4

Standard dynamics models for continuous control make use of feedforward computation to predict the conditional distribution of next state and reward given current state and action…

cs.LG202123 cited

Benchmarks for Deep Off-Policy Evaluation

Justin Fu, Mohammad Norouzi, Ofir Nachum +10

Off-policy evaluation (OPE) holds the promise of being able to leverage large, offline datasets for both evaluating and selecting complex policies for decision making. The ability…

cs.LG20214 cited

Regularized Behavior Value Estimation

Caglar Gulcehre, Sergio Gómez Colmenarejo, Ziyu Wang +7

Offline reinforcement learning restricts the learning process to rely only on logged-data without access to an environment. While this enables real-world applications, it also pose…

cs.LG202030 cited

Hyperparameter Selection for Offline Reinforcement Learning

Tom Le Paine, Cosmin Paduraru, Andrea Michi +5

Offline reinforcement learning (RL purely from logged data) is an important avenue for deploying RL techniques in real-world scenarios. However, existing hyperparameter selection m…

cs.LG2020

RL Unplugged: A Suite of Benchmarks for Offline Reinforcement Learning

Caglar Gulcehre, Ziyu Wang, Alexander Novikov +15

Offline methods for reinforcement learning have a potential to help bridge the gap between reinforcement learning research and real-world applications. They make it possible to lea…

cs.NE2019

Improving the Gating Mechanism of Recurrent Neural Networks

Albert Gu, Caglar Gulcehre, Tom Le Paine +2

Gating mechanisms are widely used in neural network models, where they allow gradients to backpropagate more easily through depth or time. However, their saturation property introd…