activity
20242026
collaborators
Showing 2024Show all

5 papers · 1 filter

eess.AS2024

Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features

Emiru Tsunoo, Yuki Saito, Wataru Nakata +1

Real-time speech enhancement (SE) is essential to online speech communication. Causal SE models use only the previous context while predicting future information, such as phoneme c…

cs.CL2024

Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens

Yosuke Kashiwagi, Hayato Futami, Emiru Tsunoo +2

In many real-world scenarios, such as meetings, multiple speakers are present with an unknown number of participants, and their utterances often overlap. We address these multi-spe…

eess.AS2024

Decoder-only Architecture for Streaming End-to-end Speech Recognition

Emiru Tsunoo, Hayato Futami, Yosuke Kashiwagi +2

Decoder-only language models (LMs) have been successfully adopted for speech-processing tasks including automatic speech recognition (ASR). The LMs have ample expressiveness and pe…

cs.SD2024

Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting

Yosuke Kashiwagi, Hayato Futami, Emiru Tsunoo +2

End-to-end multilingual speech recognition models handle multiple languages through a single model, often incorporating language identification to automatically detect the language…

cs.CL2024

Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model

Hayato Futami, Siddhant Arora, Yosuke Kashiwagi +2

Recently, multi-task spoken language understanding (SLU) models have emerged, designed to address various speech processing tasks. However, these models often rely on a large numbe…