1 citations · 1 across the 7 of their papers we have counts for
7 papers
DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
Baihan Li, Zeyu Xie, Xuenan Xu +5
Audio generation has attracted significant attention. Despite remarkable enhancement in audio quality, existing models overlook diversity evaluation. This is partially due to the l…
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
Zeyu Xie, Xuenan Xu, Zhizheng Wu +1
Recently, audio generation tasks have attracted considerable research interests. Precise temporal controllability is essential to integrate audio generation with real applications.…
AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
Zeyu Xie, Xuenan Xu, Zhizheng Wu +1
Recent advancements in audio generation have enabled the creation of high-fidelity audio clips from free-form textual descriptions. However, temporal relationships, a critical feat…
FakeSound: Deepfake General Audio Detection
Zeyu Xie, Baihan Li, Xuenan Xu +3
With the advancement of audio generation, generative models can produce highly realistic audios. However, the proliferation of deepfake general audio can pose negative consequences…
A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
Xuenan Xu, Xiaohang Xu, Zeyu Xie +3
Recently, there has been an increasing focus on audio-text cross-modal learning. However, most of the existing audio-text datasets contain only simple descriptions of sound events.…
Enhancing Audio Generation Diversity with Visual Information
Zeyu Xie, Baihan Li, Xuenan Xu +2
Audio and sound generation has garnered significant attention in recent years, with a primary focus on improving the quality of generated audios. However, there has been limited re…