activity
20202022
most citedDecision Transformer: Reinforcement Learning via Sequence Modeling

465 citations · 506 across the 7 of their papers we have counts for

collaborators

11 papers

cs.LG202212 cited

CIC: Contrastive Intrinsic Control for Unsupervised Skill Discovery

Michael Laskin, Hao Liu, Xue Bin Peng +3

We introduce Contrastive Intrinsic Control (CIC), an algorithm for unsupervised skill discovery that maximizes the mutual information between state-transitions and latent skill vec…

cs.LG202111 cited

URLB: Unsupervised Reinforcement Learning Benchmark

Michael Laskin, Denis Yarats, Hao Liu +6

Deep Reinforcement Learning (RL) has emerged as a powerful paradigm to solve a range of complex yet specific control tasks. Yet training generalist agents that can quickly adapt to…

cs.LG20217 cited

Skill Preferences: Learning to Extract and Execute Robotic Skills from Human Feedback

Xiaofei Wang, Kimin Lee, Kourosh Hakhamaneshi +2

A promising approach to solving challenging long-horizon tasks has been to extract behavior priors (skills) by fitting generative models to large offline datasets of demonstrations…

cs.LG2021465 cited

Decision Transformer: Reinforcement Learning via Sequence Modeling

Lili Chen, Kevin Lu, Aravind Rajeswaran +6

We introduce a framework that abstracts Reinforcement Learning (RL) as a sequence modeling problem. This allows us to draw upon the simplicity and scalability of the Transformer ar…

cs.LG20219 cited

Behavioral Priors and Dynamics Models: Improving Performance and Domain Transfer in Offline RL

Catherine Cang, Aravind Rajeswaran, Pieter Abbeel +1

Offline Reinforcement Learning (RL) aims to extract near-optimal policies from imperfect offline data without additional environment interactions. Extracting policies from diverse…

cs.LG20212 cited

Reinforcement Learning with Latent Flow

Wenling Shang, Xiaofei Wang, Aravind Srinivas +4

Temporal information is essential to learning effective policies with Reinforcement Learning (RL). However, current state-of-the-art RL algorithms either assume that such informati…