activity
20122022
most citedDecision Transformer: Reinforcement Learning via Sequence Modeling

465 citations · 4.1k across the 97 of their papers we have counts for

collaborators
Showing cs.LGShow all

102 papers · 1 filter

cs.LG20228 cited

Reward Uncertainty for Exploration in Preference-based Reinforcement Learning

Xinran Liang, Katherine Shu, Kimin Lee +1

Conveying complex objectives to reinforcement learning (RL) agents often requires meticulous reward engineering. Preference-based RL methods are able to learn a more flexible rewar…

cs.LG20224 cited

Chain of Thought Imitation with Procedure Cloning

Mengjiao Yang, Dale Schuurmans, Pieter Abbeel +1

Imitation learning aims to extract high-performance policies from logged demonstrations of expert behavior. It is common to frame imitation learning as a supervised learning proble…

cs.LG20227 cited

An Empirical Investigation of Representation Learning for Imitation

Xin Chen, Sam Toyer, Cody Wild +9

Imitation learning often needs a large demonstration set in order to handle the full range of situations that an agent might find itself in during deployment. However, collecting e…

cs.LG20222 cited

Imitating, Fast and Slow: Robust learning from demonstrations via decision-time planning

Carl Qi, Pieter Abbeel, Aditya Grover

The goal of imitation learning is to mimic expert behavior from demonstrations, without access to an explicit reward signal. A popular class of approach infers the (unknown) reward…

cs.LG20222 cited

Pretraining Graph Neural Networks for few-shot Analog Circuit Modeling and Design

Kourosh Hakhamaneshi, Marcel Nassar, Mariano Phielipp +2

Being able to predict the performance of circuits without running expensive simulations is a desired capability that can catalyze automated design. In this paper, we present a supe…

cs.LG202212 cited

CIC: Contrastive Intrinsic Control for Unsupervised Skill Discovery

Michael Laskin, Hao Liu, Xue Bin Peng +3

We introduce Contrastive Intrinsic Control (CIC), an algorithm for unsupervised skill discovery that maximizes the mutual information between state-transitions and latent skill vec…