activity
20162024
most citedA Comprehensive Survey on 3D Content Generation

4 citations · 18 across the 12 of their papers we have counts for

collaborators

12 papers

cs.CV20241 cited

3DBench: A Scalable 3D Benchmark and Instruction-Tuning Dataset

Junjie Zhang, Tianci Hu, Xiaoshui Huang +2

Evaluating the performance of Multi-modal Large Language Models (MLLMs), integrating both point cloud and language, presents significant challenges. The lack of a comprehensive ass…

cs.CV20241 cited

Taming Stable Diffusion for Text to 360° Panorama Image Generation

Cheng Zhang, Qianyi Wu, Camilo Cruz Gambardella +4

Generative models, e.g., Stable Diffusion, have enabled the creation of photorealistic images from text prompts. Yet, the generation of 360-degree panorama images from text remains…

cs.CV20244 cited

A Comprehensive Survey on 3D Content Generation

Jian Liu, Xiaoshui Huang, Tianyu Huang +8

Recent years have witnessed remarkable advances in artificial intelligence generated content(AIGC), with diverse input modalities, e.g., text, image, video, audio and 3D. The 3D is…

cs.CV20241 cited

THOR: Text to Human-Object Interaction Diffusion via Relation Intervention

Qianyang Wu, Ye Shi, Xiaoshui Huang +3

This paper addresses new methodologies to deal with the challenging task of generating dynamic Human-Object Interactions from textual descriptions (Text2HOI). While most existing w…

cs.CV20241 cited

NeRF-Det++: Incorporating Semantic Cues and Perspective-aware Depth Supervision for Indoor Multi-View 3D Detection

Chenxi Huang, Yuenan Hou, Weicai Ye +5

NeRF-Det has achieved impressive performance in indoor multi-view 3D detection by innovatively utilizing NeRF to enhance representation learning. Despite its notable performance, w…

cs.CV20243 cited

Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models

Dingning Liu, Xiaoshui Huang, Yuenan Hou +5

In this paper, we introduce Uni3D-LLM, a unified framework that leverages a Large Language Model (LLM) to integrate tasks of 3D perception, generation, and editing within point clo…