activity
20192022
most citedGuided Source Separation Meets a Strong ASR Backend: Hitachi/Paderborn University Joint Investigation for Dinner Party ASR

12 citations · 51 across the 26 of their papers we have counts for

collaborators

33 papers

eess.AS2022

Speech separation with large-scale self-supervised learning

Zhuo Chen, Naoyuki Kanda, Jian Wu +6

Self-supervised learning (SSL) methods such as WavLM have shown promising speech separation (SS) results in small-scale simulation-based experiments. In this work, we extend the ex…

eess.AS20221 cited

Simulating realistic speech overlaps improves multi-talker ASR

Muqiao Yang, Naoyuki Kanda, Xiaofei Wang +5

Multi-talker automatic speech recognition (ASR) has been studied to generate transcriptions of natural conversation including overlapping speech of multiple speakers. Due to the di…

eess.AS2022

Self-supervised learning with bi-label masked speech prediction for streaming multi-talker speech recognition

Zili Huang, Zhuo Chen, Naoyuki Kanda +6

Self-supervised learning (SSL), which utilizes the input data itself for representation learning, has achieved state-of-the-art results for various downstream speech tasks. However…

eess.AS20222 cited

VarArray Meets t-SOT: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition

Naoyuki Kanda, Jian Wu, Xiaofei Wang +3

This paper presents a novel streaming automatic speech recognition (ASR) framework for multi-talker overlapping speech captured by a distant microphone array with an arbitrary geom…

eess.AS20221 cited

Target Speaker Voice Activity Detection with Transformers and Its Integration with End-to-End Neural Diarization

Dongmei Wang, Xiong Xiao, Naoyuki Kanda +2

This paper describes a speaker diarization model based on target speaker voice activity detection (TS-VAD) using transformers. To overcome the original TS-VAD model's drawback of b…

cs.LG20224 cited

i-Code: An Integrative and Composable Multimodal Learning Framework

Ziyi Yang, Yuwei Fang, Chenguang Zhu +17

Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to…