activity
20182022
most citedRecent Developments on ESPnet Toolkit Boosted by Conformer

40 citations · 86 across the 17 of their papers we have counts for

collaborators

27 papers

cs.CL20225 cited

Speech-to-Speech Translation For A Real-world Unwritten Language

Peng-Jen Chen, Kevin Tran, Yilin Yang +13

We study speech-to-speech translation (S2ST) that translates speech from one language into another language and focuses on building systems to support languages without standard te…

cs.CL20222 cited

Simple and Effective Unsupervised Speech Translation

Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen +5

The amount of labeled data to train models for speech tasks is limited for most languages, however, the data scarcity is exacerbated for speech translation which requires labeled d…

cs.CL2022

Non-autoregressive Error Correction for CTC-based ASR with Phone-conditioned Masked LM

Hayato Futami, Hirofumi Inaguma, Sei Ueno +3

Connectionist temporal classification (CTC) -based models are attractive in automatic speech recognition (ASR) because of their non-autoregressive nature. To take advantage of text…

cs.CL20226 cited

Distilling the Knowledge of BERT for CTC-based ASR

Hayato Futami, Hirofumi Inaguma, Masato Mimura +2

Connectionist temporal classification (CTC) -based models are attractive because of their fast inference in automatic speech recognition (ASR). Language model (LM) integration appr…

eess.AS2022

A Study of Transducer based End-to-End ASR with ESPnet: Architecture, Auxiliary Loss and Decoding Strategies

Florian Boyer, Yusuke Shinohara, Takaaki Ishii +2

In this study, we present recent developments of models trained with the RNN-T loss in ESPnet. It involves the use of various architectures such as recently proposed Conformer, mul…

eess.AS20217 cited

A Comparative Study on Non-Autoregressive Modelings for Speech-to-Text Generation

Yosuke Higuchi, Nanxin Chen, Yuya Fujita +6

Non-autoregressive (NAR) models simultaneously generate multiple outputs in a sequence, which significantly reduces the inference speed at the cost of accuracy drop compared to aut…