9 papers
T2LDM++: A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
Wentao Qu, Qi Zhang, Chenxu Wang +5
Recent progress in Text-to-Image generation benefits from large-scale Text-Image pairs. However, the scarcity of Text-LiDAR pairs often causes over-smoothed scenes and limited cont…
Unsupervised Point Cloud Pre-Training via Contrasting and Clustering
Guofeng Mei, Xiaoshui Huang, Juan Liu +2
Annotating large-scale point clouds is highly time-consuming and often infeasible for many complex real-world tasks. Point cloud pre-training has therefore become a promising strat…
Universal 3D Shape Matching via Coarse-to-Fine Language Guidance
Qinfeng Xiao, Guofeng Mei, Bo Yang +3
Establishing dense correspondences between shapes is a crucial task in computer vision and graphics, while prior approaches depend on near-isometric assumptions and homogeneous sub…
A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
Wentao Qu, Guofeng Mei, Yang Wu +3
Text-to-LiDAR generation can customize 3D data with rich structures and diverse scenes for downstream tasks. However, the scarcity of Text-LiDAR pairs often causes insufficient tra…
Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion
Wentao Qu, Guofeng Mei, Jing Wang +3
Denoising Diffusion Probabilistic Models (DDPMs) have shown success in robust 3D object detection tasks. Existing methods often rely on the score matching from 3D boxes or pre-trai…
Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
Bin Ren, Xiaoshui Huang, Mengyuan Liu +4
Vision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challe…