6 papers
MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model
Ye Tao, Wen Wu, Chao Zhang +3
Text-guided audio editing aims to modify specific acoustic events while strictly preserving non-target content. Despite recent progress, existing approaches remain fundamentally li…
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
Zihao Zheng, Zeyu Xie, Xuenan Xu +3
While recent work in controllable text-to-audio (TTA) generation has achieved fine-grained control through timestamp conditioning, its scope remains limited by audio quality and in…
When Audio Generators Become Good Listeners: Generative Features for Understanding Tasks
Zeyu Xie, Chenxing Li, Xuenan Xu +6
This work pioneers the utilization of generative features in enhancing audio understanding. Unlike conventional discriminative features that directly optimize posterior and thus em…
UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities
Xuenan Xu, Jiahao Mei, Zihao Zheng +9
Audio generation, including speech, music and sound effects, has advanced rapidly in recent years. These tasks can be divided into two categories: time-aligned (TA) tasks, where ea…
FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
Zeyu Xie, Yaoyun Zhang, Xuenan Xu +4
The rapid development of generative audio raises ethical and security concerns stemming from forged data, making deepfake sound detection an important safeguard against the malicio…
STAR: Speech-to-Audio Generation via Representation Learning
Zeyu Xie, Xuenan Xu, Yixuan Li +2
This work presents STAR, the first end-to-end speech-to-audio generation framework, designed to enhance efficiency and address error propagation inherent in cascaded systems. Unlik…