668 citations · 1k across the 13 of their papers we have counts for
21 papers
Quasi-Bayesian Dual Instrumental Variable Regression
Ziyu Wang, Yuhao Zhou, Tongzheng Ren +1
Recent years have witnessed an upsurge of interest in employing flexible machine learning models for instrumental variable (IV) regression, but the development of uncertainty quant…
Autoregressive Dynamics Models for Offline Policy Evaluation and Optimization
Michael R. Zhang, Tom Le Paine, Ofir Nachum +4
Standard dynamics models for continuous control make use of feedforward computation to predict the conditional distribution of next state and reward given current state and action…
Benchmarks for Deep Off-Policy Evaluation
Justin Fu, Mohammad Norouzi, Ofir Nachum +10
Off-policy evaluation (OPE) holds the promise of being able to leverage large, offline datasets for both evaluating and selecting complex policies for decision making. The ability…
Regularized Behavior Value Estimation
Caglar Gulcehre, Sergio Gómez Colmenarejo, Ziyu Wang +7
Offline reinforcement learning restricts the learning process to rely only on logged-data without access to an environment. While this enables real-world applications, it also pose…
Offline Learning from Demonstrations and Unlabeled Experience
Konrad Zolna, Alexander Novikov, Ksenia Konyushkova +6
Behavior cloning (BC) is often practical for robot learning because it allows a policy to be trained offline without rewards, by supervised learning on expert demonstrations. Howev…
Hyperparameter Selection for Offline Reinforcement Learning
Tom Le Paine, Cosmin Paduraru, Andrea Michi +5
Offline reinforcement learning (RL purely from logged data) is an important avenue for deploying RL techniques in real-world scenarios. However, existing hyperparameter selection m…