3 papers
cs.SD2024
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
Yi Yuan, Dongya Jia, Xiaobin Zhuang +9
Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performan…
cs.SD2023
Text-Driven Foley Sound Generation With Latent Diffusion Model
Yi Yuan, Haohe Liu, Xubo Liu +4
Foley sound generation aims to synthesise the background sound for multimedia content. Previous models usually employ a large development set with labels as input (e.g., single num…
cs.SD2023
Latent Diffusion Model Based Foley Sound Generation System For DCASE Challenge 2023 Task 7
Yi Yuan, Haohe Liu, Xubo Liu +3
Foley sound presents the background sound for multimedia content and the generation of Foley sound involves computationally modelling sound effects with specialized techniques. In…