3 papers
eess.AS2026
FSD50K-Solo: Automated Curation of Single-Source Sound Events
Ningyuan Yang, Sile Yin, Li-Chia Yang +4
High-quality training datasets are essential for the performance of neural networks. However, the audio domain still lacks a large-scale, strongly-labeled, and single-source sound…
eess.AS2024
CATSE: A Context-Aware Framework for Causal Target Sound Extraction
Shrishail Baligar, Mikolaj Kegler, Bryce Irvin +2
Target Sound Extraction (TSE) focuses on the problem of separating sources of interest, indicated by a user's cue, from the input mixture. Most existing solutions operate in an off…
eess.AS2024
Latent CLAP Loss for Better Foley Sound Synthesis
Tornike Karchkhadze, Hassan Salami Kavaki, Mohammad Rasool Izadi +5
Foley sound generation, the art of creating audio for multimedia, has recently seen notable advancements through text-conditioned latent diffusion models. These systems use multimo…