4 papers
Sound Scene Synthesis at the DCASE 2024 Challenge
Mathieu Lagrange, Junwon Lee, Modan Tailleur +5
This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis. Recent advances in sound synthesis and generative models have enabled the creation of realistic and d…
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation
Junwon Lee, Modan Tailleur, Laurie M. Heller +5
Despite significant advancements in neural text-to-audio generation, challenges persist in controllability and evaluation. This paper addresses these issues through the Sound Scene…
Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
Modan Tailleur, Junwon Lee, Mathieu Lagrange +4
This paper explores whether considering alternative domain-specific embeddings to calculate the Fréchet Audio Distance (FAD) metric can help the FAD to correlate better with percep…
Detection of Deepfake Environmental Audio
Hafsa Ouajdi, Oussama Hadder, Modan Tailleur +2
With the ever-rising quality of deep generative models, it is increasingly important to be able to discern whether the audio data at hand have been recorded or synthesized. Althoug…