6 papers
Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI
Xinhao Liu, Jiaqi Li, Youming Deng +7
Reproducible closed-loop evaluation remains a major bottleneck in Embodied AI such as visual navigation. A promising path forward is high-fidelity simulation that combines photorea…
Thinking in 360°: Humanoid Visual Search in the Wild
Heyang Yu, Yinan Han, Xiangyu Zhang +9
Humans rely on the synergistic control of head (cephalomotor) and eye (oculomotor) to efficiently search for visual information in 360°. However, prior approaches to visual search…
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
Xinhao Liu, Jintong Li, Yicheng Jiang +6
Navigating dynamic urban environments presents significant challenges for embodied agents, requiring advanced spatial reasoning and adherence to common-sense norms. Despite progres…
Multiview Scene Graph
Juexiao Zhang, Gao Zhu, Sihang Li +4
A proper scene representation is central to the pursuit of spatial intelligence where agents can robustly reconstruct and efficiently understand 3D scenes. A scene representation i…
Self-Supervised Place Recognition by Refining Temporal and Featural Pseudo Labels from Panoramic Data
Chao Chen, Zegang Cheng, Xinhao Liu +4
Visual place recognition (VPR) using deep networks has achieved state-of-the-art performance. However, most of them require a training set with ground truth sensor poses to obtain…
SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving
Yiming Li, Sihang Li, Xinhao Liu +11
Monocular scene understanding is a foundational component of autonomous systems. Within the spectrum of monocular perception topics, one crucial and useful task for holistic 3D sce…