5 papers · 1 filter
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
Emiru Tsunoo, Yuki Saito, Wataru Nakata +1
Real-time speech enhancement (SE) is essential to online speech communication. Causal SE models use only the previous context while predicting future information, such as phoneme c…
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
Yosuke Kashiwagi, Hayato Futami, Emiru Tsunoo +2
In many real-world scenarios, such as meetings, multiple speakers are present with an unknown number of participants, and their utterances often overlap. We address these multi-spe…
Decoder-only Architecture for Streaming End-to-end Speech Recognition
Emiru Tsunoo, Hayato Futami, Yosuke Kashiwagi +2
Decoder-only language models (LMs) have been successfully adopted for speech-processing tasks including automatic speech recognition (ASR). The LMs have ample expressiveness and pe…
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
Yosuke Kashiwagi, Hayato Futami, Emiru Tsunoo +2
End-to-end multilingual speech recognition models handle multiple languages through a single model, often incorporating language identification to automatically detect the language…
Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model
Hayato Futami, Siddhant Arora, Yosuke Kashiwagi +2
Recently, multi-task spoken language understanding (SLU) models have emerged, designed to address various speech processing tasks. However, these models often rely on a large numbe…