5 papers
MMTalker: Multiresolution 3D Talking Head Synthesis with Multimodal Feature Fusion
Bin Liu, Zhixiang Xiong, Zhifen He +1
Speech-driven three-dimensional (3D) facial animation synthesis aims to build a mapping from one-dimensional (1D) speech signals to time-varying 3D facial motion signals. Current m…
SketchFaceGS: Real-Time Sketch-Driven Face Editing and Generation with Gaussian Splatting
Bo Li, Jiahao Kang, Yubo Ma +4
3D Gaussian representations have emerged as a powerful paradigm for digital head modeling, achieving photorealistic quality with real-time rendering. However, intuitive and interac…
Incomplete Multi-Label Image Recognition by Co-learning Semantic-Aware Features and Label Recovery
Zhi-Fen He, Ren-Dong Xie, Bo Li +2
Multi-label image recognition with incomplete labels is a challenging yet vital task in computer vision, which faces two fundamental challenges: learning semantic-aware features an…
Medical Scene Reconstruction and Segmentation based on 3D Gaussian Representation
Bin Liu, Wenyan Tian, Huangxin Fu +3
3D reconstruction of medical images is a key technology in medical image analysis and clinical diagnosis, providing structural visualization support for disease assessment and surg…
Semantic-Aware Representation Learning via Conditional Transport for Multi-Label Image Classification
Ren-Dong Xie, Zhi-Fen He, Bo Li +2
Multi-label image classification is a critical task in machine learning that aims to accurately assign multiple labels to a single image. While existing methods often utilize atten…