activity
20242026
most citedRetrieval-Augmented Dialogue Knowledge Aggregation for Expressive Conversational Speech Synthesis

1 citations · 2 across the 10 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling

Xueqi Wang, Zhigang Wang, Runqing Zhang +2

Multimodal dialogue retrieval aims to retrieve dialogues from multimodal dialogue banks that are similar to a target dialogue in terms of both textual semantics and acoustic conver…

cs.CL2025

Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction Learning

Rui Liu, Yuan Zhao, Zhenqi Jia

The automatic movie dubbing model generates vivid speech from given scripts, replicating a speaker's timbre from a brief timbre prompt while ensuring lip-sync with the silent video…

cs.CL2025

Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis

Zhenqi Jia, Rui Liu, Berrak Sisman +1

Conversational Speech Synthesis (CSS) aims to generate speech with natural prosody by understanding the multimodal dialogue history (MDH). The latest work predicts the accurate pro…

cs.CL20251 cited

Retrieval-Augmented Dialogue Knowledge Aggregation for Expressive Conversational Speech Synthesis

Rui Liu, Zhenqi Jia, Feilong Bao +1

Conversational speech synthesis (CSS) aims to take the current dialogue (CD) history as a reference to synthesize expressive speech that aligns with the conversational style. Unlik…

cs.CL2024

Intra- and Inter-modal Context Interaction Modeling for Conversational Speech Synthesis

Zhenqi Jia, Rui Liu

Conversational Speech Synthesis (CSS) aims to effectively take the multimodal dialogue history (MDH) to generate speech with appropriate conversational prosody for target utterance…

cs.CL2024

Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling

Rui Liu, Zhenqi Jia, Jie Yang +2

Conversational Text-to-Speech (CTTS) aims to accurately express an utterance with the appropriate style within a conversational setting, which attracts more attention nowadays. Whi…