6 papers
Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency
Brian Zhu, Momen Khalil, E Harrison +17
While reinforcement learning (RL) allows generalist robot policies to continually improve during deployment, the large model size of modern generalist policies, such as VLAs, poses…
LLM-as-a-Verifier: A General-Purpose Verification Framework
Jacky Kwok, Shulu Li, Pranav Atreya +6
Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabilities of LLMs. In this work, we identify verification, the abi…
Transitive RL: Value Learning via Divide and Conquer
Seohong Park, Aditya Oberai, Pranav Atreya +1
In this work, we present Transitive Reinforcement Learning (TRL), a new value learning algorithm based on a divide-and-conquer paradigm. TRL is designed for offline goal-conditione…
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
Pranav Atreya, Karl Pertsch, Tony Lee +29
Comprehensive, unbiased, and comparable evaluation of modern generalist policies is uniquely challenging: existing approaches for robot benchmarking typically rely on heavy standar…
AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
Zhiyuan Zhou, Pranav Atreya, You Liang Tan +2
Scalable and reproducible policy evaluation has been a long-standing challenge in robot learning. Evaluations are critical to assess progress and build better policies, but evaluat…
Autonomous Improvement of Instruction Following Skills via Foundation Models
Zhiyuan Zhou, Pranav Atreya, Abraham Lee +3
Intelligent instruction-following robots capable of improving from autonomously collected experience have the potential to transform robot learning: instead of collecting costly te…