Showing 2024 · cs.SDShow all
2 papers · 2 filters
cs.SD2024
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation
Junwon Lee, Modan Tailleur, Laurie M. Heller +5
Despite significant advancements in neural text-to-audio generation, challenges persist in controllability and evaluation. This paper addresses these issues through the Sound Scene…
cs.SD2024
Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
Modan Tailleur, Junwon Lee, Mathieu Lagrange +4
This paper explores whether considering alternative domain-specific embeddings to calculate the Fréchet Audio Distance (FAD) metric can help the FAD to correlate better with perce…