19 citations · 31 across the 21 of their papers we have counts for
9 papers · 1 filter
t-SOT FNT: Streaming Multi-talker ASR with Text-only Domain Adaptation Capability
Jian Wu, Naoyuki Kanda, Takuya Yoshioka +3
Token-level serialized output training (t-SOT) was recently proposed to address the challenge of streaming multi-talker automatic speech recognition (ASR). T-SOT effectively handle…
Bilingual Streaming ASR with Grapheme units and Auxiliary Monolingual Loss
Mohammad Soleymanpour, Mahmoud Al Ismail, Fahimeh Bahmaninezhad +2
We introduce a bilingual solution to support English as secondary locale for most primary locales in hybrid automatic speech recognition (ASR) settings. Our key developments consti…
On decoder-only architecture for speech-to-text and large language model integration
Jian Wu, Yashesh Gaur, Zhuo Chen +8
Large language models (LLMs) have achieved remarkable success in the field of natural language processing, enabling better human-computer interaction using natural language. Howeve…
Speech separation with large-scale self-supervised learning
Zhuo Chen, Naoyuki Kanda, Jian Wu +6
Self-supervised learning (SSL) methods such as WavLM have shown promising speech separation (SS) results in small-scale simulation-based experiments. In this work, we extend the ex…
Simulating realistic speech overlaps improves multi-talker ASR
Muqiao Yang, Naoyuki Kanda, Xiaofei Wang +5
Multi-talker automatic speech recognition (ASR) has been studied to generate transcriptions of natural conversation including overlapping speech of multiple speakers. Due to the di…
Self-supervised learning with bi-label masked speech prediction for streaming multi-talker speech recognition
Zili Huang, Zhuo Chen, Naoyuki Kanda +6
Self-supervised learning (SSL), which utilizes the input data itself for representation learning, has achieved state-of-the-art results for various downstream speech tasks. However…