10 papers
Continuous Speech for Improved Learning Pathological Voice Disorders
Syu-Siang Wang, Chi-Te Wang, Chih-Chung Lai +2
Goal: Numerous studies had successfully differentiated normal and abnormal voice samples. Nevertheless, further classification had rarely been attempted. This study proposes a nove…
Attention-based multi-task learning for speech-enhancement and speaker-identification in multi-speaker dialogue scenario
Chiang-Jen Peng, Yun-Ju Chan, Cheng Yu +3
Multi-task learning (MTL) and attention mechanism have been proven to effectively extract robust acoustic features for various speech-related tasks in noisy environments. In this s…
Boosting Objective Scores of a Speech Enhancement Model by MetricGAN Post-processing
Szu-Wei Fu, Chien-Feng Liao, Tsun-An Hsieh +9
The Transformer architecture has demonstrated a superior ability compared to recurrent neural networks in many different natural language processing applications. Therefore, our st…
Speech Enhancement based on Denoising Autoencoder with Multi-branched Encoders
Cheng Yu, Ryandhimas E. Zezario, Syu-Siang Wang +5
Deep learning-based models have greatly advanced the performance of speech enhancement (SE) systems. However, two problems remain unsolved, which are closely related to model gener…
MoEVC: A Mixture-of-experts Voice Conversion System with Sparse Gating Mechanism for Accelerating Online Computation
Yu-Tao Chang, Yuan-Hong Yang, Yu-Huai Peng +4
With the recent advancements of deep learning technologies, the performance of voice conversion (VC) in terms of quality and similarity has been significantly improved. However, he…
Time-Domain Multi-modal Bone/air Conducted Speech Enhancement
Cheng Yu, Kuo-Hsuan Hung, Syu-Siang Wang +3
Previous studies have proven that integrating video signals, as a complementary modality, can facilitate improved performance for speech enhancement (SE). However, video clips usua…