activity
20182024
most citedGraphRevisedIE: Multimodal Information Extraction with Graph-Revised Network

19 citations · 31 across the 21 of their papers we have counts for

collaborators
Showing eess.ASShow all

9 papers · 1 filter

eess.AS2023

t-SOT FNT: Streaming Multi-talker ASR with Text-only Domain Adaptation Capability

Jian Wu, Naoyuki Kanda, Takuya Yoshioka +3

Token-level serialized output training (t-SOT) was recently proposed to address the challenge of streaming multi-talker automatic speech recognition (ASR). T-SOT effectively handle…

eess.AS2023

Bilingual Streaming ASR with Grapheme units and Auxiliary Monolingual Loss

Mohammad Soleymanpour, Mahmoud Al Ismail, Fahimeh Bahmaninezhad +2

We introduce a bilingual solution to support English as secondary locale for most primary locales in hybrid automatic speech recognition (ASR) settings. Our key developments consti…

eess.AS2023

On decoder-only architecture for speech-to-text and large language model integration

Jian Wu, Yashesh Gaur, Zhuo Chen +8

Large language models (LLMs) have achieved remarkable success in the field of natural language processing, enabling better human-computer interaction using natural language. Howeve…

eess.AS2022

Speech separation with large-scale self-supervised learning

Zhuo Chen, Naoyuki Kanda, Jian Wu +6

Self-supervised learning (SSL) methods such as WavLM have shown promising speech separation (SS) results in small-scale simulation-based experiments. In this work, we extend the ex…

eess.AS20221 cited

Simulating realistic speech overlaps improves multi-talker ASR

Muqiao Yang, Naoyuki Kanda, Xiaofei Wang +5

Multi-talker automatic speech recognition (ASR) has been studied to generate transcriptions of natural conversation including overlapping speech of multiple speakers. Due to the di…

eess.AS2022

Self-supervised learning with bi-label masked speech prediction for streaming multi-talker speech recognition

Zili Huang, Zhuo Chen, Naoyuki Kanda +6

Self-supervised learning (SSL), which utilizes the input data itself for representation learning, has achieved state-of-the-art results for various downstream speech tasks. However…