activity
20232025
most citedMLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition

24 citations · 43 across the 12 of their papers we have counts for

collaborators
Showing cs.SDShow all

8 papers · 1 filter

cs.SD2025

DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification

Qing Wang, Jixun Yao, Zhaokai Sun +3

Being a form of biometric identification, the security of the speaker identification (SID) system is of utmost importance. To better understand the robustness of SID systems, we ai…

cs.SD2024

Optimizing Dysarthria Wake-Up Word Spotting: An End-to-End Approach for SLT 2024 LRDWWS Challenge

Shuiyun Liu, Yuxiang Kong, Pengcheng Guo +4

Speech has emerged as a widely embraced user interface across diverse applications. However, for individuals with dysarthria, the inherent variability in their speech poses signifi…

cs.SD20241 cited

ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge

He Wang, Pengcheng Guo, Yue Li +13

To promote speech processing and recognition research in driving scenarios, we build on the success of the Intelligent Cockpit Speech Recognition Challenge (ICSRC) held at ISCSLP 2…

cs.SD2024

An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge

Runduo Han, Xiaopeng Yan, Weiming Xu +6

This paper describes our audio-quality-based multi-strategy approach for the audio-visual target speaker extraction (AVTSE) task in the Multi-modal Information based Speech Process…

cs.SD202424 cited

MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition

He Wang, Pengcheng Guo, Pan Zhou +1

While automatic speech recognition (ASR) systems degrade significantly in noisy environments, audio-visual speech recognition (AVSR) systems aim to complement the audio stream with…

cs.SD2023

Automatic channel selection and spatial feature integration for multi-channel speech recognition across various array topologies

Bingshen Mu, Pengcheng Guo, Dake Guo +3

Automatic Speech Recognition (ASR) has shown remarkable progress, yet it still faces challenges in real-world distant scenarios across various array topologies each with multiple r…