67 citations · 198 across the 20 of their papers we have counts for
25 papers
LongFNT: Long-form Speech Recognition with Factorized Neural Transducer
Xun Gong, Yu Wu, Jinyu Li +4
Traditional automatic speech recognition~(ASR) systems usually focus on individual utterances, without considering long-form speech with useful historical information, which is mor…
Wespeaker: A Research and Production oriented Speaker Embedding Learning Toolkit
Hongji Wang, Chengdong Liang, Shuai Wang +5
Speaker modeling is essential for many related tasks, such as speaker recognition and speaker diarization. The dominant modeling approach is fixed-dimensional vector representation…
SJTU-AISPEECH System for VoxCeleb Speaker Recognition Challenge 2022
Zhengyang Chen, Bing Han, Xu Xiang +3
This report describes the SJTU-AISPEECH system for the Voxceleb Speaker Recognition Challenge 2022. For track1, we implemented two kinds of systems, the online system and the offli…
Layer-wise Fast Adaptation for End-to-End Multi-Accent Speech Recognition
Xun Gong, Yizhou Lu, Zhikai Zhou +1
Accent variability has posed a huge challenge to automatic speech recognition~(ASR) modeling. Although one-hot accent vector based adaptation systems are commonly used, they requir…
End-to-End Multi-speaker ASR with Independent Vector Analysis
Robin Scheibler, Wangyou Zhang, Xuankai Chang +2
We develop an end-to-end system for multi-channel, multi-speaker automatic speech recognition. We propose a frontend for joint source separation and dereverberation based on the in…
Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
Fan Yu, Shiliang Zhang, Pengcheng Guo +13
The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologie…