87 citations · 152 across the 16 of their papers we have counts for
17 papers
MIMO-DBnet: Multi-channel Input and Multiple Outputs DOA-aware Beamforming Network for Speech Separation
Yanjie Fu, Haoran Yin, Meng Ge +5
Recently, many deep learning based beamformers have been proposed for multi-channel speech separation. Nevertheless, most of them rely on extra cues known in advance, such as speak…
The ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge (ICSRC): Dataset, Tracks, Baseline and Results
Ao Zhang, Fan Yu, Kaixun Huang +7
This paper summarizes the outcomes from the ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge (ICSRC). We first address the necessity of the challenge and then introduce…
I4U System Description for NIST SRE'20 CTS Challenge
Kong Aik Lee, Tomi Kinnunen, Daniele Colibro +23
This manuscript describes the I4U submission to the 2020 NIST Speaker Recognition Evaluation (SRE'20) Conversational Telephone Speech (CTS) Challenge. The I4U's submission was resu…
Monolingual Recognizers Fusion for Code-switching Speech Recognition
Tongtong Song, Qiang Xu, Haoyu Lu +5
The bi-encoder structure has been intensively investigated in code-switching (CS) automatic speech recognition (ASR). However, most existing methods require the structures of two m…
Deep Spectro-temporal Artifacts for Detecting Synthesized Speech
Xiaohui Liu, Meng Liu, Lin Zhang +7
The Audio Deep Synthesis Detection (ADD) Challenge has been held to detect generated human-like speech. With our submitted system, this paper provides an overall assessment of trac…
VCSE: Time-Domain Visual-Contextual Speaker Extraction Network
Junjie Li, Meng Ge, Zexu Pan +2
Speaker extraction seeks to extract the target speech in a multi-talker scenario given an auxiliary reference. Such reference can be auditory, i.e., a pre-recorded speech, visual,…