4 papers
DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation
Wei Zhou, Wanyi Ning, Yinshang Guo +3
Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic co…
PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction
Wanyi Ning, Wei Zhou, Yingpeng Li +3
Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision ar…
FormalASR: End-to-End Spoken Chinese to Formal Text
Wanyi Ning, Yinshang Guo, Haitao Qian +3
Automatic speech recognition (ASR) systems are typically optimized for verbatim transcription, which preserves disfluencies, filler words, and informal spoken structures that are o…
SplitMeanFlow: Interval Splitting Consistency in Few-Step Generative Modeling
Yi Guo, Wei Wang, Zhihang Yuan +8
Generative models like Flow Matching have achieved state-of-the-art performance but are often hindered by a computationally expensive iterative sampling process. To address this, r…