activity
20172025
most citedConservative Q-Learning for Offline Reinforcement Learning

538 citations · 970 across the 21 of their papers we have counts for

collaborators
Showing cs.LGShow all

25 papers · 1 filter

cs.LG2022

Data-Driven Offline Decision-Making via Invariant Representation Learning

Han Qi, Yi Su, Aviral Kumar +1

The goal in offline data-driven decision-making is synthesize decisions that optimize a black-box utility function, using a previously-collected static dataset, with no active inte…

cs.LG20222 cited

Offline RL With Realistic Datasets: Heteroskedasticity and Support Constraints

Anikait Singh, Aviral Kumar, Quan Vuong +2

Offline reinforcement learning (RL) learns policies entirely from static datasets, thereby avoiding the challenges associated with online data collection. Practical applications of…

cs.LG20222 cited

Dual Generator Offline Reinforcement Learning

Quan Vuong, Aviral Kumar, Sergey Levine +1

In offline RL, constraining the learned policy to remain close to the data is essential to prevent the policy from outputting out-of-distribution (OOD) actions with erroneously ove…

cs.LG202217 cited

When Should We Prefer Offline Reinforcement Learning Over Behavioral Cloning?

Aviral Kumar, Joey Hong, Anikait Singh +1

Offline reinforcement learning (RL) algorithms can acquire effective policies by utilizing previously collected experience, without any online interaction. It is widely understood…

cs.LG20227 cited

Design-Bench: Benchmarks for Data-Driven Offline Model-Based Optimization

Brandon Trabucco, Xinyang Geng, Aviral Kumar +1

Black-box model-based optimization (MBO) problems, where the goal is to find a design input that maximizes an unknown objective function, are ubiquitous in a wide range of domains,…

cs.LG20216 cited

A Workflow for Offline Model-Free Robotic Reinforcement Learning

Aviral Kumar, Anikait Singh, Stephen Tian +2

Offline reinforcement learning (RL) enables learning control policies by utilizing only prior experience, without any online interaction. This can allow robots to acquire generaliz…