7 papers
GLASS: Graph and Vision-Language Assisted Semantic Shape Correspondence
Qinfeng Xiao, Guofeng Mei, Qilong Liu +5
Establishing dense correspondence across 3D shapes is crucial for fundamental downstream tasks, including texture transfer, shape interpolation, and robotic manipulation. However,…
Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
Bin Ren, Xiaoshui Huang, Mengyuan Liu +4
Vision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challe…
Fully-Geometric Cross-Attention for Point Cloud Registration
Weijie Wang, Guofeng Mei, Jian Zhang +3
Point cloud registration approaches often fail when the overlap between point clouds is low due to noisy point correspondences. This work introduces a novel cross-attention mechani…
Parameter-Efficient CLIP Adaptation for 3D Understanding via Unified Tokenization
Guofeng Mei, Bin Ren, Qinfeng Xiao +8
Vision-language models, such as CLIP, encode rich semantic knowledge through large-scale image-text pretraining. Reusing these models for 3D understanding is highly desirable, beca…
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
Jinlong Li, Cristiano Saltori, Fabio Poiesi +1
The lack of a large-scale 3D-text corpus has led recent works to distill open-vocabulary knowledge from vision-language models (VLMs). However, these methods typically rely on a si…
ZeroReg: Zero-Shot Point Cloud Registration with Foundation Models
Weijie Wang, Wenqi Ren, Guofeng Mei +5
State-of-the-art 3D point cloud registration methods rely on labeled 3D datasets for training, which limits their practical applications in real-world scenarios and often hinders g…