activity
20172022
most citedUnsupervised Speech Enhancement Based on Multichannel NMF-Informed Beamforming for Noise-Robust Automatic Speech Recognition

69 citations · 109 across the 15 of their papers we have counts for

collaborators

24 papers

cs.RO2022

Alzheimer's Dementia Detection through Spontaneous Dialogue with Proactive Robotic Listeners

Yuanchao Li, Catherine Lai, Divesh Lala +2

As the aging of society continues to accelerate, Alzheimer's Disease (AD) has received more and more attention from not only medical but also other fields, such as computer science…

cs.CL2022

Non-autoregressive Error Correction for CTC-based ASR with Phone-conditioned Masked LM

Hayato Futami, Hirofumi Inaguma, Sei Ueno +3

Connectionist temporal classification (CTC) -based models are attractive in automatic speech recognition (ASR) because of their non-autoregressive nature. To take advantage of text…

cs.CL20226 cited

Distilling the Knowledge of BERT for CTC-based ASR

Hayato Futami, Hirofumi Inaguma, Masato Mimura +2

Connectionist temporal classification (CTC) -based models are attractive because of their fast inference in automatic speech recognition (ASR). Language model (LM) integration appr…

cs.CL2021

ASR Rescoring and Confidence Estimation with ELECTRA

Hayato Futami, Hirofumi Inaguma, Masato Mimura +2

In automatic speech recognition (ASR) rescoring, the hypothesis with the fewest errors should be selected from the n-best list using a language model (LM). However, LMs are usually…

eess.AS20214 cited

Non-autoregressive End-to-end Speech Translation with Parallel Autoregressive Rescoring

Hirofumi Inaguma, Yosuke Higuchi, Kevin Duh +2

This article describes an efficient end-to-end speech translation (E2E-ST) framework based on non-autoregressive (NAR) models. End-to-end speech translation models have several adv…

eess.AS2021

VAD-free Streaming Hybrid CTC/Attention ASR for Unsegmented Recording

Hirofumi Inaguma, Tatsuya Kawahara

In this work, we propose novel decoding algorithms to enable streaming automatic speech recognition (ASR) on unsegmented long-form recordings without voice activity detection (VAD)…