3 papers
cs.SD2024
An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
Runduo Han, Xiaopeng Yan, Weiming Xu +6
This paper describes our audio-quality-based multi-strategy approach for the audio-visual target speaker extraction (AVTSE) task in the Multi-modal Information based Speech Process…
cs.CV2023
R2-Talker: Realistic Real-Time Talking Head Synthesis with Hash Grid Landmarks Encoding and Progressive Multilayer Conditioning
Zhiling Ye, LiangGuo Zhang, Dingheng Zeng +2
Dynamic NeRFs have recently garnered growing attention for 3D talking portrait synthesis. Despite advances in rendering speed and visual quality, challenges persist in enhancing ef…
cs.CV2023
PMMTalk: Speech-Driven 3D Facial Animation from Complementary Pseudo Multi-modal Features
Tianshun Han, Shengnan Gui, Yiqing Huang +9
Speech-driven 3D facial animation has improved a lot recently while most related works only utilize acoustic modality and neglect the influence of visual and textual cues, leading…