12 papers
Dynamic Prosody Prediction in LLM-based TTS for Improving Speaker Similarity
Zhenwei Mou, Liping Chen, Yajun Hu +3
Personalized text-to-speech (TTS) aims to clone the target speaker in the synthesized speech, imitating both the voice and speaking style. Current large language model (LLM)-based…
DuraMark: Duration-Embedded Watermarking in LLM-based TTS
Zhenwei Mou, Weili Jiang, Liping Chen +4
Large language model (LLM)-based text-to-speech (TTS) models have achieved remarkable voice cloning capabilities, raising concerns about potential deepfake misuse. Speech watermark…
IDMap: A Pseudo-Speaker Generator Framework Based on Speaker Identity Index to Vector Mapping
Zeyan Liu, Liping Chen, Kong Aik Lee +1
Facilitated by the speech generation framework that disentangles speech into content, speaker, and prosody, voice anonymization is accomplished by substituting the original speaker…
Pinhole Effect on Linkability and Dispersion in Speaker Anonymization
Kong Aik Lee, Zeyan Liu, Liping Chen +1
Speaker anonymization aims to conceal speaker-specific attributes in speech signals, making the anonymized speech unlinkable to the original speaker identity. Recent approaches ach…
A Study of the Removability of Speaker-Adversarial Perturbations
Liping Chen, Chenyang Guo, Kong Aik Lee +2
Recent advancements in adversarial attacks have demonstrated their effectiveness in misleading speaker recognition models, making wrong predictions about speaker identities. On the…
Investigation of perception inconsistency in speaker embedding for asynchronous voice anonymization
Rui Wang, Liping Chen, Kong Aik Lee +2
Given the speech generation framework that represents the speaker attribute with an embedding vector, asynchronous voice anonymization can be achieved by modifying the speaker embe…