activity
20182024
most citedRelation Modeling with Graph Convolutional Networks for Facial Action Unit Detection

87 citations · 152 across the 17 of their papers we have counts for

collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2024

Progressive Residual Extraction based Pre-training for Speech Representation Learning

Tianrui Wang, Jin Li, Ziyang Ma +8

Self-supervised learning (SSL) has garnered significant attention in speech processing, excelling in linguistic tasks such as speech recognition. However, jointly improving the per…

eess.AS20221 cited

MIMO-DBnet: Multi-channel Input and Multiple Outputs DOA-aware Beamforming Network for Speech Separation

Yanjie Fu, Haoran Yin, Meng Ge +5

Recently, many deep learning based beamformers have been proposed for multi-channel speech separation. Nevertheless, most of them rely on extra cues known in advance, such as speak…

eess.AS2022

I4U System Description for NIST SRE'20 CTS Challenge

Kong Aik Lee, Tomi Kinnunen, Daniele Colibro +23

This manuscript describes the I4U submission to the 2020 NIST Speaker Recognition Evaluation (SRE'20) Conversational Telephone Speech (CTS) Challenge. The I4U's submission was resu…

eess.AS20221 cited

Monolingual Recognizers Fusion for Code-switching Speech Recognition

Tongtong Song, Qiang Xu, Haoyu Lu +5

The bi-encoder structure has been intensively investigated in code-switching (CS) automatic speech recognition (ASR). However, most existing methods require the structures of two m…

eess.AS2022

L-SpEx: Localized Target Speaker Extraction

Meng Ge, Chenglin Xu, Longbiao Wang +3

Speaker extraction aims to extract the target speaker's voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction…

eess.AS20205 cited

Multi-stage Speaker Extraction with Utterance and Frame-Level Reference Signals

Meng Ge, Chenglin Xu, Longbiao Wang +3

Speaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extract…