collaborators

7 papers

cs.RO2026

Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online

Yishu Li, Xinyi Mao, Ying Yuan +3

We introduce a novel History-Aware VErifier (HAVE) to disambiguate uncertain scenarios online by leveraging past interactions. Robots frequently encounter visually ambiguous object…

cs.CV2026

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation

Hengkai Tan, Yao Feng, Xinyi Mao +5

Learning generalizable manipulation policies hinges on data, yet robot manipulation data is scarce and often entangled with specific embodiments, making both cross-task and cross-p…

cs.CV2026

Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline

Guo Chen, Lidong Lu, Yicheng Liu +17

While datasets for video understanding have scaled to hour-long durations, they typically consist of densely concatenated clips that differ from natural, unscripted daily life. To…

cs.LG2026

ManiBox: Enhancing Embodied Spatial Generalization via Scalable Simulation Data Generations

Hengkai Tan, Xuezhou Xu, Chengyang Ying +7

Embodied agents require robust spatial intelligence to execute precise real-world manipulations. However, this remains a significant challenge, as current methods often struggle to…

cs.LG2025

Vidar: Embodied Video Diffusion Model for Generalist Manipulation

Yao Feng, Hengkai Tan, Xinyi Mao +5

Scaling general-purpose manipulation to new robot embodiments remains challenging: each platform typically needs large, homogeneous demonstrations, and end-to-end pixel-to-action p…

cs.RO2025

Vidarc: Embodied Video Diffusion Model for Closed-loop Control

Yao Feng, Chendong Xiang, Xinyi Mao +7

Robotic arm manipulation in data-scarce settings is a highly challenging task due to the complex embodiment dynamics and diverse contexts. Recent video-based approaches have shown…