15 papers
MMAG: A Multi-Control Mixed Audio Generation Benchmark
Zihao Zheng, Xuenan Xu, Jiahao Mei +5
Recent audio generation systems have progressed from single-modality synthesis to generating complex acoustic scenes containing speech, music, and sound effects. Therefore, evaluat…
SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations
Xiaoyu Yang, Xuenan Xu, Wenyi Yu +10
Recent audio large language models (ALLMs) are typically built upon audio encoders trained with large amounts of supervised data. Since self-supervised learning (SSL) audio encoder…
Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation
Ye Tao, Lupeng Liu, Xuenan Xu +6
Recent unified audio generation models can support diverse tasks across speech, sound effects, and music, but most of them still focus on isolated task-level synthesis. However, re…
AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling
Yiming Ren, Xuenan Xu, Ziyang Zhang +3
Despite advances in text and visual generation, creating coherent long-form audio narratives remains challenging. Existing frameworks often exhibit limitations such as mismatched c…
FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining
Xiquan Li, Xuenan Xu, Ziyang Ma +4
Contrastively pretrained audio-language models (e.g., CLAP) excel at clip-level understanding but struggle with frame-level tasks. Existing extensions fail to exploit the varying g…
CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
Zihao Zheng, Wen Wu, Chao Zhang +2
Current Text-to-Speech (TTS) systems typically use separate models for speech-prompted and text-prompted timbre control. While unifying both control signals into a single model is…