activity
20192026
most citedTowards Online End-to-end Transformer Automatic Speech Recognition

31 citations · 55 across the 36 of their papers we have counts for

collaborators
Showing 2023Show all

9 papers · 1 filter

cs.CL2023

Phoneme-aware Encoding for Prefix-tree-based Contextual ASR

Hayato Futami, Emiru Tsunoo, Yosuke Kashiwagi +3

In speech recognition applications, it is important to recognize context-specific rare words, such as proper nouns. Tree-constrained Pointer Generator (TCPGen) has shown promise fo…

cs.CL20231 cited

UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions

Siddhant Arora, Hayato Futami, Jee-weon Jung +6

Recent studies leverage large language models with multi-tasking capabilities, using natural language prompts to guide the model's behavior and surpassing performance of task-speci…

eess.AS20231 cited

Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation

Emiru Tsunoo, Hayato Futami, Yosuke Kashiwagi +2

Collecting audio-text pairs is expensive; however, it is much easier to access text-only data. Unless using shallow fusion, end-to-end automatic speech recognition (ASR) models req…

eess.AS2023

Integration of Frame- and Label-synchronous Beam Search for Streaming Encoder-decoder Speech Recognition

Emiru Tsunoo, Hayato Futami, Yosuke Kashiwagi +2

Although frame-based models, such as CTC and transducers, have an affinity for streaming automatic speech recognition, their decoding uses no future knowledge, which could lead to…

cs.CL2023

Integrating Pretrained ASR and LM to Perform Sequence Generation for Spoken Language Understanding

Siddhant Arora, Hayato Futami, Yosuke Kashiwagi +3

There has been an increased interest in the integration of pretrained speech recognition (ASR) and language models (LM) into the SLU framework. However, prior methods often struggl…

eess.AS2023

Tensor decomposition for minimization of E2E SLU model toward on-device processing

Yosuke Kashiwagi, Siddhant Arora, Hayato Futami +6

Spoken Language Understanding (SLU) is a critical speech recognition application and is often deployed on edge devices. Consequently, on-device processing plays a significant role…