collaborators

9 papers

cs.CV2026

Gaze Target Estimation Anywhere with Concepts

Xu Cao, Houze Yang, Vipin Gunda +5

Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit…

cs.RO2026

Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training

Yaxuan Li, Zhongyi Zhou, Yefei Chen +5

Post-training is essential for turning pretrained generalist robot policies into reliable task-specific controllers, but existing human-in-the-loop pipelines remain tied to physica…

cs.RO2026

dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model

Yaxuan Li, Zhongyi Zhou, Yefei Chen +2

Evaluating robotics policies across thousands of environments and thousands of tasks is infeasible with existing approaches. This motivates the need for a new methodology for scala…

cs.RO2025

ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations

Qiyuan Zeng, Chengmeng Li, Jude St. John +5

We present ActiveUMI, a framework for a data collection system that transfers in-the-wild human demonstrations to robots capable of complex bimanual manipulation. ActiveUMI couples…

cs.RO2025

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Junjie Wen, Minjie Zhu, Yichen Zhu +8

In this paper, we present DiffusionVLA, a novel framework that seamlessly combines the autoregression model with the diffusion model for learning visuomotor policy. Central to our…

cs.RO2025

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Zhongyi Zhou, Yichen Zhu, Junjie Wen +2

Vision-language-action (VLA) models have emerged as the next generation of models in robotics. However, despite leveraging powerful pre-trained Vision-Language Models (VLMs), exist…