collaborators

17 papers

cs.LG2026

Q-Learning With World Models

Perry Dong, Yueru Jia, Chelsea Finn +1

Off-policy reinforcement learning (RL) has become increasingly sample-efficient, enabling applications such as RL fine-tuning of Vision-Language-Action models into reliable, high-p…

cs.RO2026

Improving Robotic Generalist Policies via Flow Reversal Steering

Andy Tang, William Chen, Andrew Wagenmaker +2

Generalist policies can learn a wide range of skills from diverse robot datasets. In order to solve or improve on challenging new tasks, we need a way to infer and invoke the appro…

cs.LG2026

Value Flows

Perry Dong, Chongyi Zheng, Chelsea Finn +2

While most reinforcement learning methods today flatten the distribution of future returns to a single scalar value, distributional RL methods exploit the return distribution to pr…

cs.LG2026

World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry

Yuejiang Liu, Fan Feng, Lingjing Kong +6

General-purpose world models promise scalable policy evaluation, optimization, and planning, yet achieving the required level of robustness remains challenging. Unlike policy learn…

cs.RO2026

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

Perry Dong, Kuo-Han Hung, Tian Gao +2

The ability to efficiently and reliably learn new tasks has been a foundational challenge in robotics. Vision-Language-Action (VLA) models have demonstrated strong generalization a…

cs.AI2026

Poly-EPO: Training Exploratory Reasoning Models

Ifdita Hasan Orney, Jubayer Ibn Hamid, Shreya S Ramanujam +5

Exploration is a cornerstone of learning from experience: it enables agents to find solutions to complex problems, generalize to novel ones, and scale performance with test-time co…