collaborators

9 papers

cs.CV2026

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning

Wenzheng Zeng, Siyi Jiao, Chen Gao +2

Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level video understanding and generation. In this domain, autoregres…

cs.RO2026

Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos

Danze Chen, Yanzhe Chen, Qiming Huang +3

Vision-Language-Action (VLA) models require large-scale video-action pairs, yet real teleoperation remains scarce. While generated robot videos offer a scalable alternative, existi…

cs.RO2026

ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics

Ziyu Wei, Luting Wang, Chen Gao +2

Most existing vision-language manipulation research targets rigid robotic arms, whose fixed morphology limits adaptability in cluttered or confined spaces. Soft robotic arms offer…

cs.RO2026

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform

Yu Shang, Yinzhou Tang, Yiding Ma +22

World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, ex…

cs.RO2026

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation

Yanzhe Chen, Kevin Yuchen Ma, Qi Lv +4

While Vision-Language-Action (VLA) models offer broad general capabilities, deploying them on specific hardware requires real-world adaptation to bridge the embodiment gap. Since r…

cs.CV2025

Reinforcement Learning for Large Model: A Survey

Weijia Wu, Chen Gao, Joya Chen +6

Recent advances at the intersection of reinforcement learning (RL) and visual intelligence have enabled agents that not only perceive complex visual scenes but also reason, generat…