31 citations · 32 across the 7 of their papers we have counts for
8 papers
Joint Speech Recognition and Audio Captioning
Chaitanya Narisetty, Emiru Tsunoo, Xuankai Chang +3
Speech samples recorded in both indoor and outdoor environments are often contaminated with secondary audio sources. Most end-to-end monaural speech recognition systems either remo…
Run-and-back stitch search: novel block synchronous decoding for streaming encoder-decoder ASR
Emiru Tsunoo, Chaitanya Narisetty, Michael Hentschel +2
A streaming style inference of encoder-decoder automatic speech recognition (ASR) system is important for reducing latency, which is essential for interactive use cases. To this en…
Data Augmentation Methods for End-to-end Speech Recognition on Distant-Talk Scenarios
Emiru Tsunoo, Kentaro Shibata, Chaitanya Narisetty +2
Although end-to-end automatic speech recognition (E2E ASR) has achieved great performance in tasks that have numerous paired data, it is still challenging to make E2E ASR robust ag…
Gaussian Kernelized Self-Attention for Long Sequence Data and Its Application to CTC-based Speech Recognition
Yosuke Kashiwagi, Emiru Tsunoo, Shinji Watanabe
Self-attention (SA) based models have recently achieved significant performance improvements in hybrid and end-to-end automatic speech recognition (ASR) systems owing to their flex…
Streaming Transformer ASR with Blockwise Synchronous Beam Search
Emiru Tsunoo, Yosuke Kashiwagi, Shinji Watanabe
The Transformer self-attention network has shown promising performance as an alternative to recurrent neural networks in end-to-end (E2E) automatic speech recognition (ASR) systems…
Towards Online End-to-end Transformer Automatic Speech Recognition
Emiru Tsunoo, Yosuke Kashiwagi, Toshiyuki Kumakura +1
The Transformer self-attention network has recently shown promising performance as an alternative to recurrent neural networks in end-to-end (E2E) automatic speech recognition (ASR…