5 papers · 1 filter
DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
Baihan Li, Zeyu Xie, Xuenan Xu +5
Audio generation has attracted significant attention. Despite remarkable enhancement in audio quality, existing models overlook diversity evaluation. This is partially due to the l…
FakeSound: Deepfake General Audio Detection
Zeyu Xie, Baihan Li, Xuenan Xu +3
With the advancement of audio generation, generative models can produce highly realistic audios. However, the proliferation of deepfake general audio can pose negative consequences…
A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
Xuenan Xu, Xiaohang Xu, Zeyu Xie +3
Recently, there has been an increasing focus on audio-text cross-modal learning. However, most of the existing audio-text datasets contain only simple descriptions of sound events.…
Enhancing Audio Generation Diversity with Visual Information
Zeyu Xie, Baihan Li, Xuenan Xu +2
Audio and sound generation has garnered significant attention in recent years, with a primary focus on improving the quality of generated audios. However, there has been limited re…
Towards Weakly Supervised Text-to-Audio Grounding
Xuenan Xu, Ziyang Ma, Mengyue Wu +1
Text-to-audio grounding (TAG) task aims to predict the onsets and offsets of sound events described by natural language. This task can facilitate applications such as multimodal in…