9 papers
Robot Learning from Human Videos: A Survey
Junyi Ma, Erhang Zhang, Haoran Yang +4
A critical bottleneck hindering further advancement in embodied AI and robotics is the challenge of scaling robot data. To address this, the field of learning robot manipulation sk…
OGScene3D: Incremental Open-Vocabulary 3D Gaussian Scene Graph Mapping for Scene Understanding
Siting Zhu, Ziyun Lu, Guangming Wang +5
Open-vocabulary scene understanding is crucial for robotic applications, enabling robots to comprehend complex 3D environmental contexts and supporting various downstream tasks suc…
Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation
Chang Nie, Tianchen Deng, Guangming Wang +2
While recent Vision-Language-Action (VLA) models have begun to incorporate audio, they typically treat sound as static pre-execution prompts or focus exclusively on human speech. T…
RegFormer++: An Efficient Large-Scale 3D LiDAR Point Registration Network with Projection-Aware 2D Transformer
Jiuming Liu, Guangming Wang, Zhe Liu +7
Although point cloud registration has achieved remarkable advances in object-level and indoor scenes, large-scale LiDAR registration methods has been rarely explored before. Challe…
Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
Tianchen Deng, Wenhua Wu, Kunzhen Wu +7
Visual localization has traditionally been formulated as a pair-wise pose regression problem. Existing approaches mainly estimate relative poses between two images and employ a lat…
NeRFs in Robotics: A Survey
Guangming Wang, Lei Pan, Songyou Peng +7
Detailed and realistic 3D environment representations have been a long-standing goal in the fields of computer vision and robotics. The recent emergence of neural implicit represen…