collaborators

5 papers

cs.CV2025

MR-COSMO: Visual-Text Memory Recall and Direct CrOSs-MOdal Alignment Method for Query-Driven 3D Segmentation

Chade Li, Pengju Zhang, Yihong Wu

The rapid advancement of vision-language models (VLMs) in 3D domains has accelerated research in text-query-guided point cloud processing, though existing methods underperform in p…

cs.CV2025

Unveiling Hidden Vulnerabilities in Digital Human Generation via Adversarial Attacks

Zhiying Li, Yeying Jin, Fan Shen +9

Expressive human pose and shape estimation (EHPS) is crucial for digital human generation, especially in applications like live streaming. While existing research primarily focuses…

cs.CV2025

Density-aware global-local attention network for point cloud segmentation

Chade Li, Pengju Zhang, Jiaming Zhang +1

3D point cloud segmentation has a wide range of applications in areas such as autonomous driving, augmented reality, virtual reality and digital twins. The point cloud data collect…

cs.CV2025

SuperPlace: The Renaissance of Classical Feature Aggregation for Visual Place Recognition in the Era of Foundation Models

Bingxi Liu, Pengju Zhang, Li He +5

Recent visual place recognition (VPR) approaches have leveraged foundation models (FM) and introduced novel aggregation techniques. However, these methods have failed to fully expl…

cs.CV2025

DGOcc: Depth-aware Global Query-based Network for Monocular 3D Occupancy Prediction

Xu Zhao, Pengju Zhang, Bo Liu +1

Monocular 3D occupancy prediction, aiming to predict the occupancy and semantics within interesting regions of 3D scenes from only 2D images, has garnered increasing attention rece…