activity
20202022
most citedDiscretization and Re-synthesis: an alternative method to solve the Cocktail Party Problem

11 citations · 21 across the 5 of their papers we have counts for

collaborators

6 papers

eess.AS2022

Towards Low-distortion Multi-channel Speech Enhancement: The ESPNet-SE Submission to The L3DAS22 Challenge

Yen-Ju Lu, Samuele Cornell, Xuankai Chang +5

This paper describes our submission to the L3DAS22 Challenge Task 1, which consists of speech enhancement with 3D Ambisonic microphones. The core of our approach combines Deep Neur…

eess.AS2022

Conditional Diffusion Probabilistic Model for Speech Enhancement

Yen-Ju Lu, Zhong-Qiu Wang, Shinji Watanabe +3

Speech enhancement is a critical component of many user-oriented audio applications, yet current systems still suffer from distorted and unnatural outputs. While generative models…

cs.SD202211 cited

Discretization and Re-synthesis: an alternative method to solve the Cocktail Party Problem

Jing Shi, Xuankai Chang, Tomoki Hayashi +3

Deep learning based models have significantly improved the performance of speech separation with input mixtures like the cocktail party. Prominent methods (e.g., frequency-domain a…

cs.CL20218 cited

An Exploration of Self-Supervised Pretrained Representations for End-to-End Speech Recognition

Xuankai Chang, Takashi Maekaku, Pengcheng Guo +8

Self-supervised pretraining on speech data has achieved a lot of progress. High-fidelity representation of the speech signal is learned from a lot of untranscribed data and shows p…

eess.AS20202 cited

Incorporating Broad Phonetic Information for Speech Enhancement

Yen-Ju Lu, Chien-Feng Liao, Xugang Lu +2

In noisy conditions, knowing speech contents facilitates listeners to more effectively suppress background noise components and to retrieve pure speech signals. Previous studies ha…

eess.AS2020

Boosting Objective Scores of a Speech Enhancement Model by MetricGAN Post-processing

Szu-Wei Fu, Chien-Feng Liao, Tsun-An Hsieh +9

The Transformer architecture has demonstrated a superior ability compared to recurrent neural networks in many different natural language processing applications. Therefore, our st…