activity
20182022
most citedRelation Modeling with Graph Convolutional Networks for Facial Action Unit Detection

87 citations · 152 across the 16 of their papers we have counts for

collaborators

17 papers

eess.AS20221 cited

MIMO-DBnet: Multi-channel Input and Multiple Outputs DOA-aware Beamforming Network for Speech Separation

Yanjie Fu, Haoran Yin, Meng Ge +5

Recently, many deep learning based beamformers have been proposed for multi-channel speech separation. Nevertheless, most of them rely on extra cues known in advance, such as speak…

cs.SD20221 cited

The ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge (ICSRC): Dataset, Tracks, Baseline and Results

Ao Zhang, Fan Yu, Kaixun Huang +7

This paper summarizes the outcomes from the ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge (ICSRC). We first address the necessity of the challenge and then introduce…

eess.AS2022

I4U System Description for NIST SRE'20 CTS Challenge

Kong Aik Lee, Tomi Kinnunen, Daniele Colibro +23

This manuscript describes the I4U submission to the 2020 NIST Speaker Recognition Evaluation (SRE'20) Conversational Telephone Speech (CTS) Challenge. The I4U's submission was resu…

eess.AS20221 cited

Monolingual Recognizers Fusion for Code-switching Speech Recognition

Tongtong Song, Qiang Xu, Haoyu Lu +5

The bi-encoder structure has been intensively investigated in code-switching (CS) automatic speech recognition (ASR). However, most existing methods require the structures of two m…

cs.SD20222 cited

Deep Spectro-temporal Artifacts for Detecting Synthesized Speech

Xiaohui Liu, Meng Liu, Lin Zhang +7

The Audio Deep Synthesis Detection (ADD) Challenge has been held to detect generated human-like speech. With our submitted system, this paper provides an overall assessment of trac…

cs.CV2022

VCSE: Time-Domain Visual-Contextual Speaker Extraction Network

Junjie Li, Meng Ge, Zexu Pan +2

Speaker extraction seeks to extract the target speech in a multi-talker scenario given an auxiliary reference. Such reference can be auditory, i.e., a pre-recorded speech, visual,…