activity
20182022
most citedSURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

14 citations · 40 across the 9 of their papers we have counts for

collaborators

12 papers

cs.CL2022

Few-shot Subgoal Planning with Language Models

Lajanugen Logeswaran, Yao Fu, Moontae Lee +1

Pre-trained large language models have shown successful progress in many language understanding benchmarks. This work explores the capability of these models to predict actionable…

cs.LG20221 cited

Fast Inference and Transfer of Compositional Task Structures for Few-shot Task Generalization

Sungryull Sohn, Hyunjae Woo, Jongwook Choi +4

We tackle real-world problems with complex structures beyond the pixel-based game or simulator. We formulate it as a few-shot reinforcement learning problem where a task is charact…

cs.LG2022

Learning Parameterized Task Structure for Generalization to Unseen Entities

Anthony Z. Liu, Sungryull Sohn, Mahdi Qazwini +1

Real world tasks are hierarchical and compositional. Tasks can be composed of multiple subtasks (or sub-goals) that are dependent on each other. These subtasks are defined in terms…

cs.LG202214 cited

SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Jongjin Park, Younggyo Seo, Jinwoo Shin +3

Preference-based reinforcement learning (RL) has shown potential for teaching agents to perform the target tasks without a costly, pre-defined reward function by learning the rewar…

cs.LG20228 cited

Lipschitz-constrained Unsupervised Skill Discovery

Seohong Park, Jongwook Choi, Jaekyeom Kim +2

We study the problem of unsupervised skill discovery, whose goal is to learn a set of diverse and useful skills with no external reward. There have been a number of skill discovery…

cs.LG20225 cited

Environment Generation for Zero-Shot Compositional Reinforcement Learning

Izzeddin Gur, Natasha Jaques, Yingjie Miao +4

Many real-world problems are compositional - solving them requires completing interdependent sub-tasks, either in series or in parallel, that can be represented as a dependency gra…