5 citations · 9 across the 5 of their papers we have counts for
5 papers
FoleyGen: Visually-Guided Audio Generation
Xinhao Mei, Varun Nagaraja, Gael Le Lan +4
Recent advancements in audio generation have been spurred by the evolution of large-scale deep learning models and expansive datasets. However, the task of video-to-audio (V2A) gen…
Enhance audio generation controllability through representation similarity regularization
Yangyang Shi, Gael Le Lan, Varun Nagaraja +6
This paper presents an innovative approach to enhance control over audio generation by emphasizing the alignment between audio and text representations during model training. In th…
Dual Transformer Decoder based Features Fusion Network for Automated Audio Captioning
Jianyuan Sun, Xubo Liu, Xinhao Mei +3
Automated audio captioning (AAC) which generates textual descriptions of audio content. Existing AAC models achieve good results but only use the high-dimensional representation of…
Surrey System for DCASE 2022 Task 5: Few-shot Bioacoustic Event Detection with Segment-level Metric Learning
Haohe Liu, Xubo Liu, Xinhao Mei +3
Few-shot audio event detection is a task that detects the occurrence time of a novel sound class given a few examples. In this work, we propose a system based on segment-level metr…
Segment-level Metric Learning for Few-shot Bioacoustic Event Detection
Haohe Liu, Xubo Liu, Xinhao Mei +3
Few-shot bioacoustic event detection is a task that detects the occurrence time of a novel sound given a few examples. Previous methods employ metric learning to build a latent spa…