9 papers
CRAG: Can 3D Generative Models Help 3D Assembly?
Zeyu Jiang, Sihang Li, Siqi Tan +8
Most existing 3D assembly methods treat the problem as pure pose estimation, rearranging observed parts via rigid transformations. In contrast, human assembly naturally couples str…
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
Beichen Wang, Juexiao Zhang, Shuwen Dong +2
Vision Language Models (VLMs) have recently been adopted in robotics for their capability in common sense reasoning and generalizability. Existing work has applied VLMs to generate…
Co-VisiON: Co-Visibility ReasONing on Sparse Image Sets of Indoor Scenes
Chao Chen, Nobel Dang, Juexiao Zhang +7
Humans exhibit a remarkable ability to recognize co-visibility-the 3D regions simultaneously visible in multiple images-even when these images are sparsely distributed across a com…
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Irving Fang, Juexiao Zhang, Shengbang Tong +1
One promise that Vision-Language-Action (VLA) models hold over traditional imitation learning for robotics is to leverage the broad generalization capabilities of large Vision-Lang…
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
Xinhao Liu, Jintong Li, Yicheng Jiang +6
Navigating dynamic urban environments presents significant challenges for embodied agents, requiring advanced spatial reasoning and adherence to common-sense norms. Despite progres…
URLOST: Unsupervised Representation Learning without Stationarity or Topology
Zeyu Yun, Juexiao Zhang, Yann LeCun +1
Unsupervised representation learning has seen tremendous progress. However, it is constrained by its reliance on domain specific stationarity and topology, a limitation not found i…