4 citations · 18 across the 12 of their papers we have counts for
12 papers
3DBench: A Scalable 3D Benchmark and Instruction-Tuning Dataset
Junjie Zhang, Tianci Hu, Xiaoshui Huang +2
Evaluating the performance of Multi-modal Large Language Models (MLLMs), integrating both point cloud and language, presents significant challenges. The lack of a comprehensive ass…
Taming Stable Diffusion for Text to 360° Panorama Image Generation
Cheng Zhang, Qianyi Wu, Camilo Cruz Gambardella +4
Generative models, e.g., Stable Diffusion, have enabled the creation of photorealistic images from text prompts. Yet, the generation of 360-degree panorama images from text remains…
A Comprehensive Survey on 3D Content Generation
Jian Liu, Xiaoshui Huang, Tianyu Huang +8
Recent years have witnessed remarkable advances in artificial intelligence generated content(AIGC), with diverse input modalities, e.g., text, image, video, audio and 3D. The 3D is…
THOR: Text to Human-Object Interaction Diffusion via Relation Intervention
Qianyang Wu, Ye Shi, Xiaoshui Huang +3
This paper addresses new methodologies to deal with the challenging task of generating dynamic Human-Object Interactions from textual descriptions (Text2HOI). While most existing w…
NeRF-Det++: Incorporating Semantic Cues and Perspective-aware Depth Supervision for Indoor Multi-View 3D Detection
Chenxi Huang, Yuenan Hou, Weicai Ye +5
NeRF-Det has achieved impressive performance in indoor multi-view 3D detection by innovatively utilizing NeRF to enhance representation learning. Despite its notable performance, w…
Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models
Dingning Liu, Xiaoshui Huang, Yuenan Hou +5
In this paper, we introduce Uni3D-LLM, a unified framework that leverages a Large Language Model (LLM) to integrate tasks of 3D perception, generation, and editing within point clo…