4 papers
SIGMA: Bridging Structural and Distributional Gaps for Vision Foundation Model Adaptation
Lingyu Xiong, Jinjin Shi, Xuran Xu +3
Vision Foundation Models (VFMs) have demonstrated impressive representational capabilities. However, adapting them to downstream tasks via full fine-tuning incurs prohibitive compu…
GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion Transformer
Yihong Lin, Zhaoxin Fan, Xianjia Wu +6
Speech-driven talking head generation is a critical yet challenging task with applications in augmented reality and virtual human modeling. While recent approaches using autoregres…
SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local Editing
Lingyu Xiong, Xize Cheng, Jintao Tan +7
Audio-driven talking face generation aims to synthesize video with lip movements synchronized to input audio. However, current generative techniques face challenges in preserving i…
Landmark-guided Diffusion Model for High-fidelity and Temporally Coherent Talking Head Generation
Jintao Tan, Xize Cheng, Lingyu Xiong +6
Audio-driven talking head generation is a significant and challenging task applicable to various fields such as virtual avatars, film production, and online conferences. However, t…