11 citations · 21 across the 5 of their papers we have counts for
6 papers
Towards Low-distortion Multi-channel Speech Enhancement: The ESPNet-SE Submission to The L3DAS22 Challenge
Yen-Ju Lu, Samuele Cornell, Xuankai Chang +5
This paper describes our submission to the L3DAS22 Challenge Task 1, which consists of speech enhancement with 3D Ambisonic microphones. The core of our approach combines Deep Neur…
Conditional Diffusion Probabilistic Model for Speech Enhancement
Yen-Ju Lu, Zhong-Qiu Wang, Shinji Watanabe +3
Speech enhancement is a critical component of many user-oriented audio applications, yet current systems still suffer from distorted and unnatural outputs. While generative models…
Discretization and Re-synthesis: an alternative method to solve the Cocktail Party Problem
Jing Shi, Xuankai Chang, Tomoki Hayashi +3
Deep learning based models have significantly improved the performance of speech separation with input mixtures like the cocktail party. Prominent methods (e.g., frequency-domain a…
An Exploration of Self-Supervised Pretrained Representations for End-to-End Speech Recognition
Xuankai Chang, Takashi Maekaku, Pengcheng Guo +8
Self-supervised pretraining on speech data has achieved a lot of progress. High-fidelity representation of the speech signal is learned from a lot of untranscribed data and shows p…
Incorporating Broad Phonetic Information for Speech Enhancement
Yen-Ju Lu, Chien-Feng Liao, Xugang Lu +2
In noisy conditions, knowing speech contents facilitates listeners to more effectively suppress background noise components and to retrieve pure speech signals. Previous studies ha…
Boosting Objective Scores of a Speech Enhancement Model by MetricGAN Post-processing
Szu-Wei Fu, Chien-Feng Liao, Tsun-An Hsieh +9
The Transformer architecture has demonstrated a superior ability compared to recurrent neural networks in many different natural language processing applications. Therefore, our st…