collaborators

7 papers

cs.RO2026

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation

Wei-Cheng Tseng, Gashon Hussein, Yuzhu Dong +9

Evaluating generalist robot manipulation policies in the real world is expensive, slow, and difficult to scale. Action-conditioned video world models offer a scalable alternative b…

cs.LG2026

Learning Process Rewards via Success Visitation Matching for Efficient RL

Raymond Tsao, Andrew Wagenmaker, Sergey Levine

In many modern applications of reinforcement learning (RL), the natural reward for a task of interest is inherently sparse: a reward of 0 is given everywhere except when the task i…

cs.LG2026

VIMPO: Value-Implicit Policy Optimization for LLMs

Zhewei Kang, Aosong Feng, Sergey Levine +2

Reinforcement learning with verifiable rewards has become a central tool for improving the reasoning ability of large language models, but current methods face a trade-off between…

cs.LG2026

Reversal Q-Learning

Aditya Oberai, Seohong Park, Sergey Levine

Iterative generative modeling techniques, such as flow matching, provide powerful tools to model complex behaviors for effective offline reinforcement learning (RL). In this work,…

cs.LG2026

RL Token: Bootstrapping Online RL with Vision-Language-Action Models

Charles Xu, Jost Tobias Springenberg, Michael Equi +4

Vision-language-action (VLA) models can learn to perform diverse manipulation skills "out of the box," but achieving the precision and speed that real-world tasks demand requires f…

cs.LG2025

Visual Pre-Training on Unlabeled Images using Reinforcement Learning

Dibya Ghosh, Sergey Levine

In reinforcement learning (RL), value-based algorithms learn to associate each observation with the states and rewards that are likely to be reached from it. We observe that many s…