activity
20162026
most citedDecision Transformer: Reinforcement Learning via Sequence Modeling

465 citations · 522 across the 14 of their papers we have counts for

collaborators
Showing cs.LGShow all

12 papers · 1 filter

cs.LG20241 cited

Semi-Supervised One-Shot Imitation Learning

Philipp Wu, Kourosh Hakhamaneshi, Yuqing Du +3

One-shot Imitation Learning~(OSIL) aims to imbue AI agents with the ability to learn a new task from a single demonstration. To supervise the learning, OSIL typically requires a pr…

cs.LG202212 cited

CIC: Contrastive Intrinsic Control for Unsupervised Skill Discovery

Michael Laskin, Hao Liu, Xue Bin Peng +3

We introduce Contrastive Intrinsic Control (CIC), an algorithm for unsupervised skill discovery that maximizes the mutual information between state-transitions and latent skill vec…

cs.LG20221 cited

Policy Architectures for Compositional Generalization in Control

Allan Zhou, Vikash Kumar, Chelsea Finn +1

Many tasks in control, robotics, and planning can be specified using desired goal configurations for various entities in the environment. Learning goal-conditioned policies is a na…

cs.LG2021465 cited

Decision Transformer: Reinforcement Learning via Sequence Modeling

Lili Chen, Kevin Lu, Aravind Rajeswaran +6

We introduce a framework that abstracts Reinforcement Learning (RL) as a sequence modeling problem. This allows us to draw upon the simplicity and scalability of the Transformer ar…

cs.LG20219 cited

Behavioral Priors and Dynamics Models: Improving Performance and Domain Transfer in Offline RL

Catherine Cang, Aravind Rajeswaran, Pieter Abbeel +1

Offline Reinforcement Learning (RL) aims to extract near-optimal policies from imperfect offline data without additional environment interactions. Extracting policies from diverse…

cs.LG20212 cited

Reinforcement Learning with Latent Flow

Wenling Shang, Xiaofei Wang, Aravind Srinivas +4

Temporal information is essential to learning effective policies with Reinforcement Learning (RL). However, current state-of-the-art RL algorithms either assume that such informati…