9 papers
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
Wenzheng Zeng, Siyi Jiao, Chen Gao +2
Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level video understanding and generation. In this domain, autoregres…
Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos
Danze Chen, Yanzhe Chen, Qiming Huang +3
Vision-Language-Action (VLA) models require large-scale video-action pairs, yet real teleoperation remains scarce. While generated robot videos offer a scalable alternative, existi…
ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics
Ziyu Wei, Luting Wang, Chen Gao +2
Most existing vision-language manipulation research targets rigid robotic arms, whose fixed morphology limits adaptability in cluttered or confined spaces. Soft robotic arms offer…
WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform
Yu Shang, Yinzhou Tang, Yiding Ma +22
World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, ex…
Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation
Yanzhe Chen, Kevin Yuchen Ma, Qi Lv +4
While Vision-Language-Action (VLA) models offer broad general capabilities, deploying them on specific hardware requires real-world adaptation to bridge the embodiment gap. Since r…
Reinforcement Learning for Large Model: A Survey
Weijia Wu, Chen Gao, Joya Chen +6
Recent advances at the intersection of reinforcement learning (RL) and visual intelligence have enabled agents that not only perceive complex visual scenes but also reason, generat…