3 papers
eess.AS2024
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
Wenze Ren, Kuo-Hsuan Hung, Rong Chao +3
This paper addresses the prevalent issue of incorrect speech output in audio-visual speech enhancement (AVSE) systems, which is often caused by poor video quality and mismatched tr…
eess.AS2024
MC-SEMamba: A Simple Multi-channel Extension of SEMamba
Wen-Yuan Ting, Wenze Ren, Rong Chao +3
Transformer-based models have become increasingly popular and have impacted speech-processing research owing to their exceptional performance in sequence modeling. Recently, a prom…
cs.SD2024
DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset
Jiawei Du, I-Ming Lin, I-Hsiang Chiu +6
Mainstream zero-shot TTS production systems like Voicebox and Seed-TTS achieve human parity speech by leveraging Flow-matching and Diffusion models, respectively. Unfortunately, hu…