8 papers
CRAG: Can 3D Generative Models Help 3D Assembly?
Zeyu Jiang, Sihang Li, Siqi Tan +8
Most existing 3D assembly methods treat the problem as pure pose estimation, rearranging observed parts via rigid transformations. In contrast, human assembly naturally couples str…
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Irving Fang, Juexiao Zhang, Shengbang Tong +1
One promise that Vision-Language-Action (VLA) models hold over traditional imitation learning for robotics is to leverage the broad generalization capabilities of large Vision-Lang…
Co-VisiON: Co-Visibility ReasONing on Sparse Image Sets of Indoor Scenes
Chao Chen, Nobel Dang, Juexiao Zhang +7
Humans exhibit a remarkable ability to recognize co-visibility-the 3D regions simultaneously visible in multiple images-even when these images are sparsely distributed across a com…
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
Ruixuan Zhang, Beichen Wang, Juexiao Zhang +3
The increasing availability of traffic videos functioning on a 24/7/365 time scale has the great potential of increasing the spatio-temporal coverage of traffic accidents, which wi…
Multiview Scene Graph
Juexiao Zhang, Gao Zhu, Sihang Li +4
A proper scene representation is central to the pursuit of spatial intelligence where agents can robustly reconstruct and efficiently understand 3D scenes. A scene representation i…
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
Xinhao Liu, Jintong Li, Yicheng Jiang +6
Navigating dynamic urban environments presents significant challenges for embodied agents, requiring advanced spatial reasoning and adherence to common-sense norms. Despite progres…