activity
20202022
most citedRecent Developments on ESPnet Toolkit Boosted by Conformer

40 citations · 69 across the 11 of their papers we have counts for

collaborators
Showing eess.ASShow all

11 papers · 1 filter

eess.AS2024

End-to-End Speech Recognition with Pre-trained Masked Language Model

Yosuke Higuchi, Tetsuji Ogawa, Tetsunori Kobayashi +1

We present a novel approach to end-to-end automatic speech recognition (ASR) that utilizes pre-trained masked language models (LMs) to facilitate the extraction of linguistic infor…

eess.AS20241 cited

Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems

Oswald Zink, Yosuke Higuchi, Carlos Mullov +2

Effective spoken dialog systems should facilitate natural interactions with quick and rhythmic timing, mirroring human communication patterns. To reduce response times, previous ef…

eess.AS2022

Improving non-autoregressive end-to-end speech recognition with pre-trained acoustic and language models

Keqi Deng, Zehui Yang, Shinji Watanabe +3

While Transformers have achieved promising results in end-to-end (E2E) automatic speech recognition (ASR), their autoregressive (AR) structure becomes a bottleneck for speeding up…

eess.AS20217 cited

A Comparative Study on Non-Autoregressive Modelings for Speech-to-Text Generation

Yosuke Higuchi, Nanxin Chen, Yuya Fujita +6

Non-autoregressive (NAR) models simultaneously generate multiple outputs in a sequence, which significantly reduces the inference speed at the cost of accuracy drop compared to aut…

eess.AS20211 cited

Advancing Momentum Pseudo-Labeling with Conformer and Initialization Strategy

Yosuke Higuchi, Niko Moritz, Jonathan Le Roux +1

Pseudo-labeling (PL), a semi-supervised learning (SSL) method where a seed model performs self-training using pseudo-labels generated from untranscribed speech, has been shown to e…

eess.AS20214 cited

Non-autoregressive End-to-end Speech Translation with Parallel Autoregressive Rescoring

Hirofumi Inaguma, Yosuke Higuchi, Kevin Duh +2

This article describes an efficient end-to-end speech translation (E2E-ST) framework based on non-autoregressive (NAR) models. End-to-end speech translation models have several adv…