activity
20232025
most citedUni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models

3 citations · 7 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV2025

A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation

Wentao Qu, Guofeng Mei, Yang Wu +3

Text-to-LiDAR generation can customize 3D data with rich structures and diverse scenes for downstream tasks. However, the scarcity of Text-LiDAR pairs often causes insufficient tra…

cs.CV2025

Parameter-Efficient CLIP Adaptation for 3D Understanding via Unified Tokenization

Guofeng Mei, Qinfeng Xiao, Bin Ren +7

Vision-language models, such as CLIP, encode rich semantic knowledge through large-scale image-text pretraining. Reusing these models for 3D understanding is highly desirable, beca…

cs.CV2025

PSReg: Prior-guided Sparse Mixture of Experts for Point Cloud Registration

Xiaoshui Huang, Zhou Huang, Yifan Zuo +4

The discriminative feature is crucial for point cloud registration. Recent methods improve the feature discriminative by distinguishing between non-overlapping and overlapping regi…

cs.CV2024

An End-to-End Robust Point Cloud Semantic Segmentation Network with Single-Step Conditional Diffusion Models

Wentao Qu, Jing Wang, YongShun Gong +2

Existing conditional Denoising Diffusion Probabilistic Models (DDPMs) with a Noise-Conditional Framework (NCF) remain challenging for 3D scene understanding tasks, as the complex g…

cs.CV2024★ 1 cited

3DBench: A Scalable 3D Benchmark and Instruction-Tuning Dataset

Junjie Zhang, Tianci Hu, Xiaoshui Huang +2

Evaluating the performance of Multi-modal Large Language Models (MLLMs), integrating both point cloud and language, presents significant challenges. The lack of a comprehensive ass…

cs.CV2024★ 3 cited

Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models

Dingning Liu, Xiaoshui Huang, Yuenan Hou +5

In this paper, we introduce Uni3D-LLM, a unified framework that leverages a Large Language Model (LLM) to integrate tasks of 3D perception, generation, and editing within point clo…