11 papers
MMAE: A Massive Multitask Audio Editing Benchmark
Ziyang Ma, Ruiqi Yan, Ruiyang Xu +35
We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing.…
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling
Guanrou Yang, Tian Tan, Qian Chen +12
Integrating speech understanding and generation is a pivotal step toward building unified speech models. However, the different representations required for these two tasks current…
Covo-Audio Technical Report
Wenfu Wang, Chenxing Li, Liqiang Zhang +23
In this work, we present Covo-Audio, a 7B-parameter end-to-end LALM that directly processes continuous audio inputs and generates audio outputs within a single unified architecture…
LARA-Gen: Enabling Continuous Emotion Control for Music Generation Models via Latent Affective Representation Alignment
Jiahao Mei, Xuenan Xu, Zeyu Xie +4
Recent advances in text-to-music models have enabled coherent music generation from text prompts, yet fine-grained emotional control remains unresolved. We introduce LARA-Gen, a fr…
SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents
Zeyu Xie, Chenxing Li, Qiao Jin +6
Recent audio generation models typically rely on Variational Autoencoders (VAEs) and perform generation within the VAE latent space. Although VAEs excel at compression and reconstr…
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
Zihao Zheng, Zeyu Xie, Xuenan Xu +3
While recent work in controllable text-to-audio (TTA) generation has achieved fine-grained control through timestamp conditioning, its scope remains limited by audio quality and in…