5 papers
T2LDM++: A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
Wentao Qu, Qi Zhang, Chenxu Wang +5
Recent progress in Text-to-Image generation benefits from large-scale Text-Image pairs. However, the scarcity of Text-LiDAR pairs often causes over-smoothed scenes and limited cont…
A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
Wentao Qu, Guofeng Mei, Yang Wu +3
Text-to-LiDAR generation can customize 3D data with rich structures and diverse scenes for downstream tasks. However, the scarcity of Text-LiDAR pairs often causes insufficient tra…
Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
Bin Ren, Xiaoshui Huang, Mengyuan Liu +4
Vision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challe…
Parameter-Efficient CLIP Adaptation for 3D Understanding via Unified Tokenization
Guofeng Mei, Bin Ren, Qinfeng Xiao +8
Vision-language models, such as CLIP, encode rich semantic knowledge through large-scale image-text pretraining. Reusing these models for 3D understanding is highly desirable, beca…
ZeroReg: Zero-Shot Point Cloud Registration with Foundation Models
Weijie Wang, Wenqi Ren, Guofeng Mei +5
State-of-the-art 3D point cloud registration methods rely on labeled 3D datasets for training, which limits their practical applications in real-world scenarios and often hinders g…