3 papers
cs.SD2024
Joint Multi-scale Cross-lingual Speaking Style Transfer with Bidirectional Attention Mechanism for Automatic Dubbing
Jingbei Li, Sipan Li, Ping Chen +7
Automatic dubbing, which generates a corresponding version of the input speech in another language, could be widely utilized in many real-world scenarios such as video and game loc…
eess.AS2024
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
Zhichao Wang, Liumeng Xue, Qiuqiang Kong +4
Zero-shot voice conversion (VC) converts source speech into the voice of any desired speaker using only one utterance of the speaker without requiring additional model updates. Typ…
cs.SD2024
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
Haohe Liu, Yi Yuan, Xubo Liu +7
Although audio generation shares commonalities across different types of audio, such as speech, music, and sound effects, designing models for each type requires careful considerat…