10 papers · 1 filter
Decoupled Q-Chunking
Qiyang Li, Seohong Park, Sergey Levine
Temporal-difference (TD) methods learn state and action values efficiently by bootstrapping from their own future value predictions, but such a self-bootstrapping mechanism is pron…
Scalable Offline Model-Based RL with Action Chunks
Kwanyoung Park, Seohong Park, Youngwoon Lee +1
In this paper, we study whether model-based reinforcement learning (RL), in particular model-based value expansion, can provide a scalable recipe for tackling complex, long-horizon…
Transitive RL: Value Learning via Divide and Conquer
Seohong Park, Aditya Oberai, Pranav Atreya +1
In this work, we present Transitive Reinforcement Learning (TRL), a new value learning algorithm based on a divide-and-conquer paradigm. TRL is designed for offline goal-conditione…
Dual Goal Representations
Seohong Park, Deepinder Mann, Sergey Levine
In this work, we introduce dual goal representations for goal-conditioned reinforcement learning (GCRL). A dual goal representation characterizes a state by "the set of temporal di…
Horizon Reduction Makes RL Scalable
Seohong Park, Kevin Frans, Deepinder Mann +3
In this work, we study the scalability of offline reinforcement learning (RL) algorithms. In principle, a truly scalable offline RL algorithm should be able to solve any given prob…
Diffusion Guidance Is a Controllable Policy Improvement Operator
Kevin Frans, Seohong Park, Pieter Abbeel +1
At the core of reinforcement learning is the idea of learning beyond the performance in the data. However, scaling such systems has proven notoriously tricky. In contrast, techniqu…