3 papers
cs.MM2025
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Lei Zhao, Linfeng Feng, Dongxu Ge +5
With the rise of diffusion models, audio-video generation has been revolutionized. However, most existing methods rely on separate modules for each modality, with limited explorati…
cs.SD2025
DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model
Lei Zhao, Sizhou Chen, Linfeng Feng +4
Text-to-audio (TTA), which generates audio signals from textual descriptions, has received huge attention in recent years. However, recent works focused on text to monaural audio o…
eess.AS2025
AudioSpa: Spatializing Sound Events with Text
Linfeng Feng, Lei Zhao, Boyu Zhu +2
Text-to-audio (TTA) systems have recently demonstrated strong performance in synthesizing monaural audio from text. However, the task of generating binaural spatial audio from text…