2 citations · 2 across the 4 of their papers we have counts for
4 papers
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
Ruibo Fu, Xin Qi, Zhengqi Wen +10
Speaker adaptation, which involves cloning voices from unseen speakers in the Text-to-Speech task, has garnered significant interest due to its numerous applications in multi-media…
Optimization of Autonomous Driving Image Detection Based on RFAConv and Triplet Attention
Zhipeng Ling, Qi Xin, Yiyu Lin +2
YOLOv8 plays a crucial role in the realm of autonomous driving, owing to its high-speed target detection, precise identification and positioning, and versatile compatibility across…
A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge
Xiaopeng Wang, Yi Lu, Xin Qi +4
This paper presents the development of a speech synthesis system for the LIMMITS'24 Challenge, focusing primarily on Track 2. The objective of the challenge is to establish a multi…
MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation
Ruibo Fu, Shuchen Shi, Hongming Guo +12
Foley audio, critical for enhancing the immersive experience in multimedia content, faces significant challenges in the AI-generated content (AIGC) landscape. Despite advancements…