activity
20242026
collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2025

Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis

Zhenqi Jia, Rui Liu, Berrak Sisman +1

Conversational Speech Synthesis (CSS) aims to generate speech with natural prosody by understanding the multimodal dialogue history (MDH). The latest work predicts the accurate pro…

cs.CL2025

NE-PADD: Leveraging Named Entity Knowledge for Robust Partial Audio Deepfake Detection via Attention Aggregation

Huhong Xian, Rui Liu, Berrak Sisman +1

Different from traditional sentence-level audio deepfake detection (ADD), partial audio deepfake detection (PADD) requires frame-level positioning of the location of fake speech. W…

cs.CL2025

Retrieval-Augmented Dialogue Knowledge Aggregation for Expressive Conversational Speech Synthesis

Rui Liu, Zhenqi Jia, Feilong Bao +1

Conversational speech synthesis (CSS) aims to take the current dialogue (CD) history as a reference to synthesize expressive speech that aligns with the conversational style. Unlik…

cs.CL2024

Intra- and Inter-modal Context Interaction Modeling for Conversational Speech Synthesis

Zhenqi Jia, Rui Liu

Conversational Speech Synthesis (CSS) aims to effectively take the multimodal dialogue history (MDH) to generate speech with appropriate conversational prosody for target utterance…

cs.CL2024

FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency

Rui Liu, Jiatian Xi, Ziyue Jiang +1

Text-based speech editing (TSE) allows users to edit speech by modifying the corresponding text directly without altering the original recording. Current TSE techniques often focus…

cs.CL2024

Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling

Rui Liu, Zhenqi Jia, Jie Yang +2

Conversational Text-to-Speech (CTTS) aims to accurately express an utterance with the appropriate style within a conversational setting, which attracts more attention nowadays. Whi…