most citedMIMO-DoAnet: Multi-channel Input and Multiple Outputs DoA Network with Unknown Number of Sound Sources

3 citations · 3 across the 6 of their papers we have counts for

collaborators

7 papers

cs.MM2024

AIMDiT: Modality Augmentation and Interaction via Multimodal Dimension Transformation for Emotion Recognition in Conversations

Sheng Wu, Jiaxing Liu, Longbiao Wang +3

Emotion Recognition in Conversations (ERC) is a popular task in natural language processing, which aims to recognize the emotional state of the speaker in conversations. While curr…

cs.SD2023

Rethinking the visual cues in audio-visual speaker extraction

Junjie Li, Meng Ge, Zexu pan +4

The Audio-Visual Speaker Extraction (AVSE) algorithm employs parallel video recording to leverage two visual cues, namely speaker identity and synchronization, to enhance performan…

eess.AS2023

Locate and Beamform: Two-dimensional Locating All-neural Beamformer for Multi-channel Speech Separation

Yanjie Fu, Meng Ge, Honglong Wang +7

Recently, stunning improvements on multi-channel speech separation have been achieved by neural beamformers when direction information is available. However, most of them neglect t…

cs.SD2023

speech and noise dual-stream spectrogram refine network with speech distortion loss for robust speech recognition

Haoyu Lu, Nan Li, Tongtong Song +4

In recent years, the joint training of speech enhancement front-end and automatic speech recognition (ASR) back-end has been widely used to improve the robustness of ASR systems. T…

cs.SD2023

Time-domain Speech Enhancement Assisted by Multi-resolution Frequency Encoder and Decoder

Hao Shi, Masato Mimura, Longbiao Wang +2

Time-domain speech enhancement (SE) has recently been intensively investigated. Among recent works, DEMUCS introduces multi-resolution STFT loss to enhance performance. However, so…

cs.CL2022

Language-specific Characteristic Assistance for Code-switching Speech Recognition

Tongtong Song, Qiang Xu, Meng Ge +5

Dual-encoder structure successfully utilizes two language-specific encoders (LSEs) for code-switching speech recognition. Because LSEs are initialized by two pre-trained language-s…