4 papers
What Limits Vision-and-Language Navigation ?
Yunheng Wang, Yuetong Fang, Taowen Wang +9
Vision-and-Language Navigation (VLN) is a cornerstone of embodied intelligence. However, current agents often suffer from significant performance degradation when transitioning fro…
Morphology-Consistent Humanoid Interaction through Robot-Centric Video Synthesis
Weisheng Xu, Jian Li, Yi Gu +12
Equipping humanoid robots with versatile interaction skills typically requires either extensive policy training or explicit human-to-robot motion retargeting. However, learning-bas…
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
Jiaxi Zhang, Yunheng Wang, Wei Lu +8
3D Visual Grounding (3DVG) localizes objects from natural-language descriptions in 3D scenes and is fundamental to embodied AI applications. Although foundation models enable open-…
SeMv-3D: Towards Concurrency of Semantic and Multi-view Consistency in General Text-to-3D Generation
Xiao Cai, Pengpeng Zeng, Lianli Gao +3
General Text-to-3D (GT23D) generation is crucial for creating diverse 3D content across objects and scenes, yet it faces two key challenges: 1) ensuring semantic consistency betwee…