activity
20192022
most citedTowards Online End-to-end Transformer Automatic Speech Recognition

31 citations · 32 across the 7 of their papers we have counts for

collaborators

8 papers

eess.AS2022

Joint Speech Recognition and Audio Captioning

Chaitanya Narisetty, Emiru Tsunoo, Xuankai Chang +3

Speech samples recorded in both indoor and outdoor environments are often contaminated with secondary audio sources. Most end-to-end monaural speech recognition systems either remo…

eess.AS2022

Run-and-back stitch search: novel block synchronous decoding for streaming encoder-decoder ASR

Emiru Tsunoo, Chaitanya Narisetty, Michael Hentschel +2

A streaming style inference of encoder-decoder automatic speech recognition (ASR) system is important for reducing latency, which is essential for interactive use cases. To this en…

eess.AS2021

Data Augmentation Methods for End-to-end Speech Recognition on Distant-Talk Scenarios

Emiru Tsunoo, Kentaro Shibata, Chaitanya Narisetty +2

Although end-to-end automatic speech recognition (E2E ASR) has achieved great performance in tasks that have numerous paired data, it is still challenging to make E2E ASR robust ag…

eess.AS2021

Gaussian Kernelized Self-Attention for Long Sequence Data and Its Application to CTC-based Speech Recognition

Yosuke Kashiwagi, Emiru Tsunoo, Shinji Watanabe

Self-attention (SA) based models have recently achieved significant performance improvements in hybrid and end-to-end automatic speech recognition (ASR) systems owing to their flex…

eess.AS2020

Streaming Transformer ASR with Blockwise Synchronous Beam Search

Emiru Tsunoo, Yosuke Kashiwagi, Shinji Watanabe

The Transformer self-attention network has shown promising performance as an alternative to recurrent neural networks in end-to-end (E2E) automatic speech recognition (ASR) systems…

eess.AS201931 cited

Towards Online End-to-end Transformer Automatic Speech Recognition

Emiru Tsunoo, Yosuke Kashiwagi, Toshiyuki Kumakura +1

The Transformer self-attention network has recently shown promising performance as an alternative to recurrent neural networks in end-to-end (E2E) automatic speech recognition (ASR…