activity
20192026
most citedTowards Online End-to-end Transformer Automatic Speech Recognition

31 citations · 36 across the 29 of their papers we have counts for

collaborators
Showing eess.ASShow all

13 papers · 1 filter

eess.AS2025

Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting

Emiru Tsunoo, Hayato Futami, Yosuke Kashiwagi +2

For streaming speech recognition, a Transformer-based encoder has been widely used with block processing. Although many studies addressed improving emission latency of transducers,…

eess.AS2024

Decoder-only Architecture for Streaming End-to-end Speech Recognition

Emiru Tsunoo, Hayato Futami, Yosuke Kashiwagi +2

Decoder-only language models (LMs) have been successfully adopted for speech-processing tasks including automatic speech recognition (ASR). The LMs have ample expressiveness and pe…

eess.AS2023★ 1 cited

Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation

Emiru Tsunoo, Hayato Futami, Yosuke Kashiwagi +2

Collecting audio-text pairs is expensive; however, it is much easier to access text-only data. Unless using shallow fusion, end-to-end automatic speech recognition (ASR) models req…

eess.AS2023

Integration of Frame- and Label-synchronous Beam Search for Streaming Encoder-decoder Speech Recognition

Emiru Tsunoo, Hayato Futami, Yosuke Kashiwagi +2

Although frame-based models, such as CTC and transducers, have an affinity for streaming automatic speech recognition, their decoding uses no future knowledge, which could lead to…

eess.AS2023

Tensor decomposition for minimization of E2E SLU model toward on-device processing

Yosuke Kashiwagi, Siddhant Arora, Hayato Futami +6

Spoken Language Understanding (SLU) is a critical speech recognition application and is often deployed on edge devices. Consequently, on-device processing plays a significant role…

eess.AS2022

Joint Speech Recognition and Audio Captioning

Chaitanya Narisetty, Emiru Tsunoo, Xuankai Chang +3

Speech samples recorded in both indoor and outdoor environments are often contaminated with secondary audio sources. Most end-to-end monaural speech recognition systems either remo…