activity
20182025
most citedRecent Developments on ESPnet Toolkit Boosted by Conformer

40 citations · 224 across the 26 of their papers we have counts for

collaborators
Showing 2020 · eess.ASShow all

10 papers · 2 filters

eess.AS2020★ 6 cited

The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans

Shinji Watanabe, Florian Boyer, Xuankai Chang +12

This paper describes the recent development of ESPnet (https://github.com/espnet/espnet), an end-to-end speech processing toolkit. This project was initiated in December 2017 to ma…

eess.AS2020

ESPnet-se: end-to-end speech enhancement and separation toolkit designed for asr integration

Chenda Li, Jing Shi, Wangyou Zhang +8

We present ESPnet-SE, which is designed for the quick development of speech enhancement and speech separation systems in a single framework, along with the optional downstream spee…

eess.AS2020★ 40 cited

Recent Developments on ESPnet Toolkit Boosted by Conformer

Pengcheng Guo, Florian Boyer, Xuankai Chang +12

In this study, we present recent developments on ESPnet: End-to-End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-…

eess.AS2020

Any-to-One Sequence-to-Sequence Voice Conversion using Self-Supervised Discrete Speech Representations

Wen-Chin Huang, Yi-Chiao Wu, Tomoki Hayashi +1

We present a novel approach to any-to-one (A2O) voice conversion (VC) in a sequence-to-sequence (seq2seq) framework. A2O VC aims to convert any speaker, including those unseen duri…

eess.AS2020★ 7 cited

The Sequence-to-Sequence Baseline for the Voice Conversion Challenge 2020: Cascading ASR and TTS

Wen-Chin Huang, Tomoki Hayashi, Shinji Watanabe +1

This paper presents the sequence-to-sequence (seq2seq) baseline system for the voice conversion challenge (VCC) 2020. We consider a naive approach for voice conversion (VC), which…

eess.AS2020★ 3 cited

Pretraining Techniques for Sequence-to-Sequence Voice Conversion

Wen-Chin Huang, Tomoki Hayashi, Yi-Chiao Wu +2

Sequence-to-sequence (seq2seq) voice conversion (VC) models are attractive owing to their ability to convert prosody. Nonetheless, without sufficient data, seq2seq VC models can su…