activity
20172022
most citedDecision Transformer: Reinforcement Learning via Sequence Modeling

465 citations · 740 across the 17 of their papers we have counts for

collaborators

25 papers

cs.CV20221 cited

HARP: Autoregressive Latent Video Prediction with High-Fidelity Image Generator

Younggyo Seo, Kimin Lee, Fangchen Liu +2

Video prediction is an important yet challenging problem; burdened with the tasks of generating future frames and learning environment dynamics. Recently, autoregressive latent vid…

cs.LG20228 cited

Reward Uncertainty for Exploration in Preference-based Reinforcement Learning

Xinran Liang, Katherine Shu, Kimin Lee +1

Conveying complex objectives to reinforcement learning (RL) agents often requires meticulous reward engineering. Preference-based RL methods are able to learn a more flexible rewar…

cs.LG202214 cited

SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Jongjin Park, Younggyo Seo, Jinwoo Shin +3

Preference-based reinforcement learning (RL) has shown potential for teaching agents to perform the target tasks without a costly, pre-defined reward function by learning the rewar…

cs.LG20213 cited

B-Pref: Benchmarking Preference-Based Reinforcement Learning

Kimin Lee, Laura Smith, Anca Dragan +1

Reinforcement learning (RL) requires access to a reward function that incentivizes the right behavior, but these are notoriously hard to specify for complex tasks. Preference-based…

cs.LG202111 cited

URLB: Unsupervised Reinforcement Learning Benchmark

Michael Laskin, Denis Yarats, Hao Liu +6

Deep Reinforcement Learning (RL) has emerged as a powerful paradigm to solve a range of complex yet specific control tasks. Yet training generalist agents that can quickly adapt to…

cs.LG20217 cited

Skill Preferences: Learning to Extract and Execute Robotic Skills from Human Feedback

Xiaofei Wang, Kimin Lee, Kourosh Hakhamaneshi +2

A promising approach to solving challenging long-horizon tasks has been to extract behavior priors (skills) by fitting generative models to large offline datasets of demonstrations…