6 papers
UniVerse: Benchmarking and Enhancing LALMs on Culturally Inclusive Low-Resource Music Understanding
Ziya Zhou, Shangda Wu, Shenyang Xu +16
Recent advances in large audio-language models (LALMs) have significantly improved performance in tasks such as music captioning, genre classification, and sound event detection. H…
UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling
Chuanbo Zhu, Wuyou Zhou, Rongxiu Zhong +4
Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-level content modification and ty…
SqueezeComposer: Temporal Speed-up is A Simple Trick for Long-form Music Composing
Jianyi Chen, Rongxiu Zhong, Shilei Zhang +4
Composing coherent long-form music remains a significant challenge due to the complexity of modeling long-range dependencies and the prohibitive memory and computational requiremen…
Residual Speaker Representation for One-Shot Voice Conversion
Le Xu, Jiangyan Yi, Tao Wang +4
Recently, there have been significant advancements in voice conversion, resulting in high-quality performance. However, there are still two critical challenges in this field. First…
Latent linguistic embedding for cross-lingual text-to-speech and voice conversion
Hieu-Thi Luong, Junichi Yamagishi
As the recently proposed voice cloning system, NAUTILUS, is capable of cloning unseen voices using untranscribed speech, we investigate the feasibility of using it to develop a uni…
Towards Fine-Grained Prosody Control for Voice Conversion
Zheng Lian, Zhengqi Wen
In a typical voice conversion system, prior works utilize various acoustic features (e.g., the pitch, voiced/unvoiced flag, aperiodicity) of the source speech to control the prosod…