2 citations · 5 across the 5 of their papers we have counts for
5 papers
LLMs Meet Multimodal Generation and Editing: A Survey
Yingqing He, Zhaoyang Liu, Jingye Chen +13
With the recent advancement in large language models (LLMs), there is a growing interest in combining LLMs with multimodal learning. Previous surveys of multimodal large language m…
CoMoSVC: Consistency Model-based Singing Voice Conversion
Yiwen Lu, Zhen Ye, Wei Xue +3
The diffusion-based Singing Voice Conversion (SVC) methods have achieved remarkable performances, producing natural audios with high similarity to the target timbre. However, the i…
Continual Learning with Dirichlet Generative-based Rehearsal
Min Zeng, Wei Xue, Qifeng Liu +1
Recent advancements in data-driven task-oriented dialogue systems (ToDs) struggle with incremental learning due to computational constraints and time-consuming issues. Continual Le…
NAS-FM: Neural Architecture Search for Tunable and Interpretable Sound Synthesis based on Frequency Modulation
Zhen Ye, Wei Xue, Xu Tan +2
Developing digital sound synthesizers is crucial to the music industry as it provides a low-cost way to produce high-quality sounds with rich timbres. Existing traditional synthesi…
Learn to Sing by Listening: Building Controllable Virtual Singer by Unsupervised Learning from Voice Recordings
Wei Xue, Yiwen Wang, Qifeng Liu +1
The virtual world is being established in which digital humans are created indistinguishable from real humans. Producing their audio-related capabilities is crucial since voice con…