41 citations · 122 across the 9 of their papers we have counts for
11 papers
Discretization and Re-synthesis: an alternative method to solve the Cocktail Party Problem
Jing Shi, Xuankai Chang, Tomoki Hayashi +3
Deep learning based models have significantly improved the performance of speech separation with input mixtures like the cocktail party. Prominent methods (e.g., frequency-domain a…
Closing the Gap Between Time-Domain Multi-Channel Speech Enhancement on Real and Simulation Conditions
Wangyou Zhang, Jing Shi, Chenda Li +2
The deep learning based time-domain models, e.g. Conv-TasNet, have shown great potential in both single-channel and multi-channel speech enhancement. However, many experiments on t…
An Exploration of Self-Supervised Pretrained Representations for End-to-End Speech Recognition
Xuankai Chang, Takashi Maekaku, Pengcheng Guo +8
Self-supervised pretraining on speech data has achieved a lot of progress. High-fidelity representation of the speech signal is learned from a lot of untranscribed data and shows p…
The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans
Shinji Watanabe, Florian Boyer, Xuankai Chang +12
This paper describes the recent development of ESPnet (https://github.com/espnet/espnet), an end-to-end speech processing toolkit. This project was initiated in December 2017 to ma…
Audio-visual Speech Separation with Adversarially Disentangled Visual Representation
Peng Zhang, Jiaming Xu, Jing shi +2
Speech separation aims to separate individual voice from an audio mixture of multiple simultaneous talkers. Although audio-only approaches achieve satisfactory performance, they bu…
Recent Developments on ESPnet Toolkit Boosted by Conformer
Pengcheng Guo, Florian Boyer, Xuankai Chang +12
In this study, we present recent developments on ESPnet: End-to-End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-…