most citedConservative Q-Learning for Offline Reinforcement Learning

538 citations · 812 across the 6 of their papers we have counts for

collaborators

6 papers

cs.LG2020538 cited

Conservative Q-Learning for Offline Reinforcement Learning

Aviral Kumar, Aurick Zhou, George Tucker +1

Effectively leveraging large, previously collected datasets in reinforcement learning (RL) is a key challenge for large-scale real-world applications. Offline RL algorithms promise…

cs.LG201916 cited

Reward-Conditioned Policies

Aviral Kumar, Xue Bin Peng, Sergey Levine

Reinforcement learning offers the promise of automating the acquisition of complex behavioral skills. However, compared to commonly used and well-understood supervised learning met…

cs.LG201919 cited

Model Inversion Networks for Model-Based Optimization

Aviral Kumar, Sergey Levine

In this work, we aim to solve data-driven optimization problems, where the goal is to find an input that maximizes an unknown score function given access to a dataset of inputs wit…

cs.LG2019166 cited

Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Xue Bin Peng, Aviral Kumar, Grace Zhang +1

In this paper, we aim to develop a simple and scalable reinforcement learning algorithm that uses standard supervised learning methods as subroutines. Our goal is an algorithm that…

cs.LG201940 cited

Calibration of Encoder Decoder Models for Neural Machine Translation

Aviral Kumar, Sunita Sarawagi

We study the calibration of several state of the art neural machine translation(NMT) systems built on attention-based encoder-decoder models. For structured outputs like in NMT, ca…

cs.LG201933 cited

Diagnosing Bottlenecks in Deep Q-learning Algorithms

Justin Fu, Aviral Kumar, Matthew Soh +1

Q-learning methods represent a commonly used class of algorithms in reinforcement learning: they are generally efficient and simple, and can be combined readily with function appro…