4 papers
EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
Li Zhou, Hao Jiang, Junjie Li +2
Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-awa…
Do We Really Need GNNs with Explicit Structural Modeling? MLPs Suffice for Language Model Representations
Li Zhou, Hao Jiang, Junjie Li +4
Explicit structural information has been proven to be encoded by Graph Neural Networks (GNNs), serving as auxiliary knowledge to enhance model capabilities and improve performance…
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
Junjie Li, Ke Zhang, Shuai Wang +3
Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world sce…
On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
Junjie Li, Ke Zhang, Shuai Wang +3
Deep learning technologies have significantly advanced the performance of target speaker extraction (TSE) tasks. To enhance the generalization and robustness of these algorithms wh…