most citedHow Bad Are Artifacts?: Analyzing the Impact of Speech Enhancement Errors on ASR

5 citations · 5 across the 5 of their papers we have counts for

collaborators

5 papers

eess.AS2024

Alignment-Free Training for Transducer-based Multi-Talker ASR

Takafumi Moriya, Shota Horiguchi, Marc Delcroix +5

Extending the RNN Transducer (RNNT) to recognize multi-talker speech is essential for wider automatic speech recognition (ASR) applications. Multi-talker RNNT (MT-RNNT) aims to ach…

eess.AS2024

NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge

Naoyuki Kamo, Naohiro Tawara, Atsushi Ando +15

We present a distant automatic speech recognition (DASR) system developed for the CHiME-8 DASR track. It consists of a diarization first pipeline. For diarization, we use end-to-en…

eess.AS2024

Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance

Tsubasa Ochiai, Kazuma Iwamoto, Marc Delcroix +4

It is challenging to improve automatic speech recognition (ASR) performance in noisy conditions with a single-channel speech enhancement (SE) front-end. This is generally attribute…

cs.SD2024

Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters

Kenichi Fujita, Hiroshi Sato, Takanori Ashihara +4

The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce sp…

eess.AS20225 cited

How Bad Are Artifacts?: Analyzing the Impact of Speech Enhancement Errors on ASR

Kazuma Iwamoto, Tsubasa Ochiai, Marc Delcroix +4

It is challenging to improve automatic speech recognition (ASR) performance in noisy conditions with single-channel speech enhancement (SE). In this paper, we investigate the cause…