activity
20172021
most citedHyperparameter Selection for Offline Reinforcement Learning

30 citations · 84 across the 5 of their papers we have counts for

collaborators

9 papers

cs.LG202123 cited

Benchmarks for Deep Off-Policy Evaluation

Justin Fu, Mohammad Norouzi, Ofir Nachum +10

Off-policy evaluation (OPE) holds the promise of being able to leverage large, offline datasets for both evaluating and selecting complex policies for decision making. The ability…

math.OC2021

Automatic differentiation for Riemannian optimization on low-rank matrix and tensor-train manifolds

Alexander Novikov, Maxim Rakhuba, Ivan Oseledets

In scientific computing and machine learning applications, matrices and more general multidimensional arrays (tensors) can often be approximated with the help of low-rank decomposi…

cs.LG20207 cited

Semi-supervised reward learning for offline reinforcement learning

Ksenia Konyushkova, Konrad Zolna, Yusuf Aytar +4

In offline reinforcement learning (RL) agents are trained using a logged dataset. It appears to be the most natural route to attack real-life applications because in domains such a…

cs.LG202014 cited

Offline Learning from Demonstrations and Unlabeled Experience

Konrad Zolna, Alexander Novikov, Ksenia Konyushkova +6

Behavior cloning (BC) is often practical for robot learning because it allows a policy to be trained offline without rewards, by supervised learning on expert demonstrations. Howev…

cs.LG202030 cited

Hyperparameter Selection for Offline Reinforcement Learning

Tom Le Paine, Cosmin Paduraru, Andrea Michi +5

Offline reinforcement learning (RL purely from logged data) is an important avenue for deploying RL techniques in real-world scenarios. However, existing hyperparameter selection m…

cs.LG2020

RL Unplugged: A Suite of Benchmarks for Offline Reinforcement Learning

Caglar Gulcehre, Ziyu Wang, Alexander Novikov +15

Offline methods for reinforcement learning have a potential to help bridge the gap between reinforcement learning research and real-world applications. They make it possible to lea…