4 papers
Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex Recordings
He Zhao, Hangting Chen, Jianwei Yu +1
Target speaker extraction (TSE) aims to extract the target speaker's voice from the input mixture. Previous studies have concentrated on high-overlapping scenarios. However, real-w…
Consistent and Relevant: Rethink the Query Embedding in General Sound Separation
Yuanyuan Wang, Hangting Chen, Dongchao Yang +4
The query-based audio separation usually employs specific queries to extract target sources from a mixture of audio signals. Currently, most query-based separation models need addi…
AutoPrep: An Automatic Preprocessing Framework for In-the-Wild Speech Data
Jianwei Yu, Hangting Chen, Yanyao Bian +6
Recently, the utilization of extensive open-sourced text data has significantly advanced the performance of text-based large language models (LLMs). However, the use of in-the-wild…
Complexity Scaling for Speech Denoising
Hangting Chen, Jianwei Yu, Chao Weng
Computational complexity is critical when deploying deep learning-based speech denoising models for on-device applications. Most prior research focused on optimizing model architec…