5 citations
- City University of Hong KongHK4 papers
- Peking UniversityCN2 papers
- Tongji UniversityCN2 papers
- University of Technology SydneyAU2 papers
- Westlake UniversityCN2 papers
- Xiamen UniversityCN2 papers
- Centre National de la Recherche ScientifiqueFR1 paper
- China Institute of Finance and Capital MarketsCN1 paper
- Chinese University of Hong KongHK1 paper
- Dalian Ocean UniversityCN1 paper
- Dalian UniversityCN1 paper
- Dalian University of TechnologyCN1 paper
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026★ 5 cited
ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance
Yongkang Cheng, Mingjiang Liang, Shaoli Huang +3
Existing gesture generation methods primarily focus on upper body gestures based on audio features, neglecting speech content, emotion, and locomotion. These limitations result in…
cs.SD2026
Leveraging large multimodal models for audio-video deepfake detection: a pilot study
Songjun Cao, Yuqi Li, Yunpeng Luo +2
Audio-visual deepfake detection (AVD) is increasingly important as modern generators can fabricate convincing speech and video. Most current multimodal detectors are small, task-sp…