4 papers
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
Junming Huang, Chi Wang, Letian Li +5
Large Language Models(LLMs) have revolutionized text generation and multimodal perception,but their capabilities in 3D content generation remain underexplored. Existing methods com…
HetScene: Heterogeneity-Aware Diffusion for Dense Indoor Scene Generation
Zini Chen, Junming Huang, Rong Zhang +4
Generating controllable and physically plausible indoor scenes is a pivotal prerequisite for constructing high-fidelity simulation environments for embodied AI. However, existing d…
VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
Donglin Huang, Yongyuan Li, Tianhang Liu +4
Existing for audio- and pose-driven human animation methods often struggle with stiff head movements and blurry hands, primarily due to the weak correlation between audio and head…
BuildingBlock: A Hybrid Approach for Structured Building Generation
Junming Huang, Chi Wang, Letian Li +3
Three-dimensional building generation is vital for applications in gaming, virtual reality, and digital twins, yet current methods face challenges in producing diverse, structured,…