activity
20242026
collaborators

6 papers

cs.SD2026

MeMo: Attentional Momentum for Real-Time Audio-Visual Target Speaker Extraction Under Impaired Visual Conditions

Junjie Li, Wenxuan Wu, Shuai Wang +4

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate a target speaker's voice from multi-speaker environments by leveraging visual cues as guidance. However, the perform…

cs.CL2026

MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

Dingdong Wang, Junan Li, Jincenzi Wu +4

Speech inherently contains rich acoustic information that extends far beyond the textual language. In real-world spoken language understanding, effective interpretation often requi…

cs.SD2025

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction

Wenxuan Wu, Shuai Wang, Xixin Wu +2

Audio-visual target speaker extraction (AV-TSE) models primarily rely on target visual cues to isolate the target speaker's voice from others. We know that humans leverage linguist…

cs.SD2025

AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction

Wenxuan Wu, Xueyuan Chen, Shuai Wang +5

Audio-Visual Target Speaker Extraction (AV-TSE) aims to mimic the human ability to enhance auditory perception using visual cues. Although numerous models have been proposed recent…

cs.SD2025

UniSep: Universal Target Audio Separation with Language Models at Scale

Yuanyuan Wang, Hangting Chen, Dongchao Yang +7

We propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio. Distinguished from previous studies, UniSep…

cs.CL2024

A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models

Dingdong Wang, Mingyu Cui, Dongchao Yang +2

With the rise of Speech Large Language Models (Speech LLMs), there has been growing interest in discrete speech tokens for their ability to integrate with text-based tokens seamles…