7 papers
TC-IDM: Grounding Video Generation for Executable Zero-shot Robot Motion
Weishi Mi, Yong Bao, Xiaowei Chi +7
The vision-language-action (VLA) paradigm has enabled powerful robotic control by leveraging vision-language models, but its reliance on large-scale, high-quality robot data limits…
Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test
Chun-Kai Fan, Xiaowei Chi, Xiaozhu Ju +18
As world models gain momentum in Embodied AI, an increasing number of works explore using video foundation models as predictive world models for downstream embodied tasks like 3D p…
WoW: Towards a World omniscient World model Through Embodied Interaction
Xiaowei Chi, Peidong Jia, Chun-Kai Fan +33
Humans develop an understanding of intuitive physics through active interaction with the world. This approach is in stark contrast to current video models, such as Sora, which rely…
FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning
Jiajun Cao, Qizhe Zhang, Peidong Jia +11
Vision-Language-Action (VLA) models have demonstrated significant potential in complex scene understanding and action reasoning, leading to their increasing adoption in end-to-end…
FastInit: Fast Noise Initialization for Temporally Consistent Video Generation
Chengyu Bai, Yuming Li, Zhongyu Zhao +5
Video generation has made significant strides with the development of diffusion models; however, achieving high temporal consistency remains a challenging task. Recently, FreeInit…
MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis
Xiaowei Chi, Kuangzhi Ge, Jiaming Liu +9
Video Generation Models (VGMs) have become powerful backbones for Vision-Language-Action (VLA) models, leveraging large-scale pretraining for robust dynamics modeling. However, cur…