5 papers
MeMo: Attentional Momentum for Real-Time Audio-Visual Target Speaker Extraction Under Impaired Visual Conditions
Junjie Li, Wenxuan Wu, Shuai Wang +4
Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate a target speaker's voice from multi-speaker environments by leveraging visual cues as guidance. However, the perform…
EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
Li Zhou, Hao Jiang, Junjie Li +2
Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-awa…
Do We Really Need GNNs with Explicit Structural Modeling? MLPs Suffice for Language Model Representations
Li Zhou, Hao Jiang, Junjie Li +4
Explicit structural information has been proven to be encoded by Graph Neural Networks (GNNs), serving as auxiliary knowledge to enhance model capabilities and improve performance…
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
Junjie Li, Ke Zhang, Shuai Wang +3
Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world sce…
Multi-Level Speaker Representation for Target Speaker Extraction
Ke Zhang, Junjie Li, Shuai Wang +4
Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the refere…