Recent Developments on ESPnet Toolkit Boosted by Conformer
arXiv:2010.13956
Abstract
In this study, we present recent developments on ESPnet: End-to-End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-augmented Transformer. This paper shows the results for a wide range of end-to-end speech processing applications, such as automatic speech recognition (ASR), speech translations (ST), speech separation (SS) and text-to-speech (TTS). Our experiments reveal various training tips and significant performance benefits obtained with the Conformer on different tasks. These results are competitive or even outperform the current state-of-art Transformer models. We are preparing to release all-in-one recipes using open source and publicly available corpora for all the above tasks with pre-trained models. Our aim for this work is to contribute to our research community by reducing the burden of preparing state-of-the-art research environments usually requiring high resources.
References in corpus (1)
Cited by in corpus (6)
- Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition
- Leveraging End-to-End ASR for Endangered Language Documentation: An Empirical Study on Yoloxóchitl Mixtec
- The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans
- Efficient End-to-End Speech Recognition Using Performers in Conformers
- Tiny Transducer: A Highly-efficient Speech Recognition Model on Edge Devices
- Advanced Long-context End-to-end Speech Recognition Using Context-expanded Transformers