8 citations · 15 across the 6 of their papers we have counts for
8 papers · 1 filter
Conditional Diffusion Probabilistic Model for Speech Enhancement
Yen-Ju Lu, Zhong-Qiu Wang, Shinji Watanabe +3
Speech enhancement is a critical component of many user-oriented audio applications, yet current systems still suffer from distorted and unnatural outputs. While generative models…
Attention-based multi-task learning for speech-enhancement and speaker-identification in multi-speaker dialogue scenario
Chiang-Jen Peng, Yun-Ju Chan, Cheng Yu +3
Multi-task learning (MTL) and attention mechanism have been proven to effectively extract robust acoustic features for various speech-related tasks in noisy environments. In this s…
HLT-NUS Submission for NIST 2019 Multimedia Speaker Recognition Evaluation
Rohan Kumar Das, Ruijie Tao, Jichen Yang +3
This work describes the speaker verification system developed by Human Language Technology Laboratory, National University of Singapore (HLT-NUS) for 2019 NIST Multimedia Speaker R…
Waveform-based Voice Activity Detection Exploiting Fully Convolutional networks with Multi-Branched Encoders
Cheng Yu, Kuo-Hsuan Hung, I-Fan Lin +3
In this study, we propose an encoder-decoder structured system with fully convolutional networks to implement voice activity detection (VAD) directly on the time-domain waveform. T…
Boosting Objective Scores of a Speech Enhancement Model by MetricGAN Post-processing
Szu-Wei Fu, Chien-Feng Liao, Tsun-An Hsieh +9
The Transformer architecture has demonstrated a superior ability compared to recurrent neural networks in many different natural language processing applications. Therefore, our st…
Speech Enhancement based on Denoising Autoencoder with Multi-branched Encoders
Cheng Yu, Ryandhimas E. Zezario, Syu-Siang Wang +5
Deep learning-based models have greatly advanced the performance of speech enhancement (SE) systems. However, two problems remain unsolved, which are closely related to model gener…