activity
20242026
collaborators

5 papers

cs.SD2026

MeMo: Attentional Momentum for Real-Time Audio-Visual Target Speaker Extraction Under Impaired Visual Conditions

Junjie Li, Wenxuan Wu, Shuai Wang +4

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate a target speaker's voice from multi-speaker environments by leveraging visual cues as guidance. However, the perform…

eess.AS2026

EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis

Li Zhou, Hao Jiang, Junjie Li +2

Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-awa…

cs.CL2025

Do We Really Need GNNs with Explicit Structural Modeling? MLPs Suffice for Language Model Representations

Li Zhou, Hao Jiang, Junjie Li +4

Explicit structural information has been proven to be encoded by Graph Neural Networks (GNNs), serving as auxiliary knowledge to enhance model capabilities and improve performance…

cs.SD2025

MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues

Junjie Li, Ke Zhang, Shuai Wang +3

Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world sce…

eess.AS2024

Multi-Level Speaker Representation for Target Speaker Extraction

Ke Zhang, Junjie Li, Shuai Wang +4

Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the refere…