activity
20232026
most citedMLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition

24 citations · 26 across the 5 of their papers we have counts for

collaborators

7 papers

eess.AS2026

MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios

Zhaokai Sun, Shuai Wang, Zhennan Lin +6

Spoken Language Understanding (SLU) is moving from task-specific pipelines toward large audio language models (LALMs) that generate natural-language responses. However, existing sp…

cs.SD2025

Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM

Zhaokai Sun, Li Zhang, Qing Wang +2

Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work prop…

eess.AS20241 cited

The NPU-ASLP-LiAuto System Description for Visual Speech Recognition in CNVSRC 2023

He Wang, Pengcheng Guo, Wei Chen +2

This paper delineates the visual speech recognition (VSR) system introduced by the NPU-ASLP-LiAuto (Team 237) in the first Chinese Continuous Visual Speech Recognition Challenge (C…

cs.SD20241 cited

ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge

He Wang, Pengcheng Guo, Yue Li +13

To promote speech processing and recognition research in driving scenarios, we build on the success of the Intelligent Cockpit Speech Recognition Challenge (ICSRC) held at ISCSLP 2…

cs.SD202424 cited

MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition

He Wang, Pengcheng Guo, Pan Zhou +1

While automatic speech recognition (ASR) systems degrade significantly in noisy environments, audio-visual speech recognition (AVSR) systems aim to complement the audio stream with…

eess.AS2023

U2-KWS: Unified Two-pass Open-vocabulary Keyword Spotting with Keyword Bias

Ao Zhang, Pan Zhou, Kaixun Huang +3

Open-vocabulary keyword spotting (KWS), which allows users to customize keywords, has attracted increasingly more interest. However, existing methods based on acoustic models and p…