6 papers
SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution
Yiren Song, Yihan Wang, Xiyao Deng +2
Visual prediction has emerged as a promising paradigm for embodied control, where future observations are generated and then translated into actions. However, dense video generatio…
OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation
Yiren Song, Xiyao Deng, Pei Yang +2
Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling scalable data generation for e…
NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions
Haolin Yang, Yuxing Long, Zhuoyuan Yu +8
Instruction-following navigation is a key step toward embodied intelligence. Prior benchmarks mainly focus on semantic understanding but overlook systematically evaluating navigati…
Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
Yi Yang, Xueqi Li, Yiyang Chen +7
Recent advances in Vision-Language-Action (VLA) models demonstrate that visual signals can effectively complement sparse action supervisions. However, letting VLA directly predict…
ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
Congzhi Zhang, Zhibin Wang, Yinchao Ma +5
While Reinforcement Learning with Verifiable Reward (RLVR) significantly advances image reasoning in Large Vision-Language Models (LVLMs), its application to complex video reasonin…
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
Yi Yang, Jiaxuan Sun, Siqi Kou +2
Real-world embodied agents face long-horizon tasks, characterized by high-level goals demanding multi-step solutions beyond single actions. Successfully navigating these requires b…