most citedLLMs Meet Multimodal Generation and Editing: A Survey

2 citations · 5 across the 5 of their papers we have counts for

collaborators

5 papers

cs.AI20242 cited

LLMs Meet Multimodal Generation and Editing: A Survey

Yingqing He, Zhaoyang Liu, Jingye Chen +13

With the recent advancement in large language models (LLMs), there is a growing interest in combining LLMs with multimodal learning. Previous surveys of multimodal large language m…

eess.AS2024

CoMoSVC: Consistency Model-based Singing Voice Conversion

Yiwen Lu, Zhen Ye, Wei Xue +3

The diffusion-based Singing Voice Conversion (SVC) methods have achieved remarkable performances, producing natural audios with high similarity to the target timbre. However, the i…

cs.CL20231 cited

Continual Learning with Dirichlet Generative-based Rehearsal

Min Zeng, Wei Xue, Qifeng Liu +1

Recent advancements in data-driven task-oriented dialogue systems (ToDs) struggle with incremental learning due to computational constraints and time-consuming issues. Continual Le…

cs.SD20231 cited

NAS-FM: Neural Architecture Search for Tunable and Interpretable Sound Synthesis based on Frequency Modulation

Zhen Ye, Wei Xue, Xu Tan +2

Developing digital sound synthesizers is crucial to the music industry as it provides a low-cost way to produce high-quality sounds with rich timbres. Existing traditional synthesi…

cs.SD20231 cited

Learn to Sing by Listening: Building Controllable Virtual Singer by Unsupervised Learning from Voice Recordings

Wei Xue, Yiwen Wang, Qifeng Liu +1

The virtual world is being established in which digital humans are created indistinguishable from real humans. Producing their audio-related capabilities is crucial since voice con…