3 papers
cs.SD2026
Hierarchical Codec Diffusion for Video-to-Speech Generation
Jiaxin Ye, Gaoxiang Cong, Chenhui Wang +4
Video-to-Speech (VTS) generation aims to synthesize speech from a silent video without auditory signals. However, existing VTS methods disregard the hierarchical nature of speech,…
cs.SD2025
Cross-attention and Self-attention for Audio-visual Speaker Diarization in MISP-Meeting Challenge
Zhaoyang Li, Haodong Zhou, Longjie Luo +4
This paper presents the system developed for Task 1 of the Multi-modal Information-based Speech Processing (MISP) 2025 Challenge. We introduce CASA-Net, an embedding fusion method…
cs.SD2025
Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm
Zhaoyang Li, Jie Wang, XiaoXiao Li +4
In speaker diarization, traditional clustering-based methods remain widely used in real-world applications. However, these methods struggle with the complex distribution of speaker…