9 papers
Xemo-Talker: Unlock Emotions Explicitly for Audio-Driven Talking Portrait Synthesis
Chaolong Yang, Yinuo Guo, Kai Yao +6
Precise emotion control in audio-driven talking heads remains a challenge due to the reliance on implicit emotion regulation in existing systems, which often leads to indirect and…
Unlock Pose Diversity: Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking Portrait
Chaolong Yang, Kai Yao, Yuyao Yan +7
Audio-driven single-image talking portrait generation plays a crucial role in virtual reality, digital human creation, and filmmaking. Existing approaches are generally categorized…
BFANet: Revisiting 3D Semantic Segmentation with Boundary Feature Analysis
Weiguang Zhao, Rui Zhang, Qiufeng Wang +2
3D semantic segmentation plays a fundamental and crucial role to understand 3D scenes. While contemporary state-of-the-art techniques predominantly concentrate on elevating the ove…
Consistency Diffusion Models for Single-Image 3D Reconstruction with Priors
Chenru Jiang, Chengrui Zhang, Xi Yang +4
This paper delves into the study of 3D point cloud reconstruction from a single image. Our objective is to develop the Consistency Diffusion Model, exploring synergistic 2D and 3D…
HOMER: Homography-Based Efficient Multi-view 3D Object Removal
Jingcheng Ni, Weiguang Zhao, Daniel Wang +4
3D object removal is an important sub-task in 3D scene editing, with broad applications in scene understanding, augmented reality, and robotics. However, existing methods struggle…
Towards Training-Free Open-World Classification with 3D Generative Models
Xinzhe Xia, Weiguang Zhao, Yuyao Yan +4
3D open-world classification is a challenging yet essential task in dynamic and unstructured real-world scenarios, requiring both open-category and open-pose recognition. To addres…