activity
20162026
most citedOpen X-Embodiment: Robotic Learning Datasets and RT-X Models

103 citations · 969 across the 139 of their papers we have counts for

collaborators
Showing cs.LGShow all

36 papers · 1 filter

cs.LG2026

Q-Learning With World Models

Perry Dong, Yueru Jia, Chelsea Finn +1

Off-policy reinforcement learning (RL) has become increasingly sample-efficient, enabling applications such as RL fine-tuning of Vision-Language-Action models into reliable, high-p…

cs.LG2026

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

Perry Dong, Ron Polonsky, Dorsa Sadigh +1

Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question:…

cs.LG2026

FASTER: Value-Guided Sampling for Fast RL

Perry Dong, Alexander Swerdlow, Dorsa Sadigh +1

Some of the most performant reinforcement learning algorithms today can be prohibitively expensive as they use test-time scaling methods such as sampling multiple action candidates…

cs.LG2025

Value Flows

Perry Dong, Chongyi Zheng, Chelsea Finn +2

While most reinforcement learning methods today flatten the distribution of future returns to a single scalar value, distributional RL methods exploit the return distribution to pr…

cs.LG2025

Polychromic Objectives for Reinforcement Learning

Jubayer Ibn Hamid, Ifdita Hasan Orney, Ellen Xu +2

Reinforcement learning fine-tuning (RLFT) is a dominant paradigm for improving pretrained policies for downstream tasks. These pretrained policies, trained on large datasets, produ…

cs.LG2025

EXPO: Stable Reinforcement Learning with Expressive Policies

Perry Dong, Qiyang Li, Dorsa Sadigh +1

We study the problem of training and fine-tuning expressive policies with online reinforcement learning (RL) given an offline dataset. Training expressive policy classes with onlin…