4 papers
GARF: Learning Generalizable 3D Reassembly for Real-World Fractures
Sihang Li, Zeyu Jiang, Grace Chen +9
3D reassembly is a challenging spatial intelligence task with broad applications across scientific domains. While large-scale synthetic datasets have fueled promising learning-base…
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
Beichen Wang, Juexiao Zhang, Shuwen Dong +2
Vision Language Models (VLMs) have recently been adopted in robotics for their capability in common sense reasoning and generalizability. Existing work has applied VLMs to generate…
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Irving Fang, Juexiao Zhang, Shengbang Tong +1
One promise that Vision-Language-Action (VLA) models hold over traditional imitation learning for robotics is to leverage the broad generalization capabilities of large Vision-Lang…
FusionSense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction
Irving Fang, Kairui Shi, Xujin He +7
Humans effortlessly integrate common-sense knowledge with sensory input from vision and touch to understand their surroundings. Emulating this capability, we introduce FusionSense,…