31 citations · 51 across the 9 of their papers we have counts for
10 papers · 1 filter
Joint Speech Recognition and Audio Captioning
Chaitanya Narisetty, Emiru Tsunoo, Xuankai Chang +3
Speech samples recorded in both indoor and outdoor environments are often contaminated with secondary audio sources. Most end-to-end monaural speech recognition systems either remo…
Run-and-back stitch search: novel block synchronous decoding for streaming encoder-decoder ASR
Emiru Tsunoo, Chaitanya Narisetty, Michael Hentschel +2
A streaming style inference of encoder-decoder automatic speech recognition (ASR) system is important for reducing latency, which is essential for interactive use cases. To this en…
Polyphone disambiguation and accent prediction using pre-trained language models in Japanese TTS front-end
Rem Hida, Masaki Hamada, Chie Kamada +3
Although end-to-end text-to-speech (TTS) models can generate natural speech, challenges still remain when it comes to estimating sentence-level phonetic and prosodic information fr…
Ensemble of ACCDOA- and EINV2-based Systems with D3Nets and Impulse Response Simulation for Sound Event Localization and Detection
Kazuki Shimada, Naoya Takahashi, Yuichiro Koyama +4
This report describes our systems submitted to the DCASE2021 challenge task 3: sound event localization and detection (SELD) with directional interference. Our previous system base…
Data Augmentation Methods for End-to-end Speech Recognition on Distant-Talk Scenarios
Emiru Tsunoo, Kentaro Shibata, Chaitanya Narisetty +2
Although end-to-end automatic speech recognition (E2E ASR) has achieved great performance in tasks that have numerous paired data, it is still challenging to make E2E ASR robust ag…
Gaussian Kernelized Self-Attention for Long Sequence Data and Its Application to CTC-based Speech Recognition
Yosuke Kashiwagi, Emiru Tsunoo, Shinji Watanabe
Self-attention (SA) based models have recently achieved significant performance improvements in hybrid and end-to-end automatic speech recognition (ASR) systems owing to their flex…