3 papers
cs.CV2025
MR-COSMO: Visual-Text Memory Recall and Direct CrOSs-MOdal Alignment Method for Query-Driven 3D Segmentation
Chade Li, Pengju Zhang, Yihong Wu
The rapid advancement of vision-language models (VLMs) in 3D domains has accelerated research in text-query-guided point cloud processing, though existing methods underperform in p…
cs.CV2025
Density-aware global-local attention network for point cloud segmentation
Chade Li, Pengju Zhang, Jiaming Zhang +1
3D point cloud segmentation has a wide range of applications in areas such as autonomous driving, augmented reality, virtual reality and digital twins. The point cloud data collect…
cs.CV2025
DGOcc: Depth-aware Global Query-based Network for Monocular 3D Occupancy Prediction
Xu Zhao, Pengju Zhang, Bo Liu +1
Monocular 3D occupancy prediction, aiming to predict the occupancy and semantics within interesting regions of 3D scenes from only 2D images, has garnered increasing attention rece…