3 papers
cs.SD2025
Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis
Zhitong Zhou, Qingqing Zhang, Lei Luo +2
Full-duplex, spontaneous conversational data are essential for enhancing the naturalness and interactivity of synthesized speech in conversational TTS systems. We present two open-…
eess.AS2023
Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
Jinzuomu Zhong, Yang Li, Hui Huang +6
In expressive and controllable Text-to-Speech (TTS), explicit prosodic features significantly improve the naturalness and controllability of synthesised speech. However, manual pro…
cs.SD2023
TranssionADD: A multi-frame reinforcement based sequence tagging model for audio deepfake detection
Jie Liu, Zhiba Su, Hui Huang +5
Thanks to recent advancements in end-to-end speech modeling technology, it has become increasingly feasible to imitate and clone a user`s voice. This leads to a significant challen…