activity
20182022
most citedNeural Speaker Diarization with Speaker-Wise Chain Rule

41 citations · 122 across the 9 of their papers we have counts for

collaborators

11 papers

cs.SD202211 cited

Discretization and Re-synthesis: an alternative method to solve the Cocktail Party Problem

Jing Shi, Xuankai Chang, Tomoki Hayashi +3

Deep learning based models have significantly improved the performance of speech separation with input mixtures like the cocktail party. Prominent methods (e.g., frequency-domain a…

eess.AS20215 cited

Closing the Gap Between Time-Domain Multi-Channel Speech Enhancement on Real and Simulation Conditions

Wangyou Zhang, Jing Shi, Chenda Li +2

The deep learning based time-domain models, e.g. Conv-TasNet, have shown great potential in both single-channel and multi-channel speech enhancement. However, many experiments on t…

cs.CL20218 cited

An Exploration of Self-Supervised Pretrained Representations for End-to-End Speech Recognition

Xuankai Chang, Takashi Maekaku, Pengcheng Guo +8

Self-supervised pretraining on speech data has achieved a lot of progress. High-fidelity representation of the speech signal is learned from a lot of untranscribed data and shows p…

eess.AS20206 cited

The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans

Shinji Watanabe, Florian Boyer, Xuankai Chang +12

This paper describes the recent development of ESPnet (https://github.com/espnet/espnet), an end-to-end speech processing toolkit. This project was initiated in December 2017 to ma…

cs.SD20203 cited

Audio-visual Speech Separation with Adversarially Disentangled Visual Representation

Peng Zhang, Jiaming Xu, Jing shi +2

Speech separation aims to separate individual voice from an audio mixture of multiple simultaneous talkers. Although audio-only approaches achieve satisfactory performance, they bu…

eess.AS202040 cited

Recent Developments on ESPnet Toolkit Boosted by Conformer

Pengcheng Guo, Florian Boyer, Xuankai Chang +12

In this study, we present recent developments on ESPnet: End-to-End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-…