4 citations · 12 across the 8 of their papers we have counts for
7 papers · 1 filter
FSD50K-Solo: Automated Curation of Single-Source Sound Events
Ningyuan Yang, Sile Yin, Li-Chia Yang +4
High-quality training datasets are essential for the performance of neural networks. However, the audio domain still lacks a large-scale, strongly-labeled, and single-source sound…
PAS-SE: Personalized Auxiliary-Sensor Speech Enhancement for Voice Pickup in Hearables
Mattes Ohlenbusch, Mikolaj Kegler, Marko Stamenovic
Speech enhancement for voice pickup in hearables aims to improve the user's voice by suppressing noise and interfering talkers, while maintaining own-voice quality. For single-chan…
CATSE: A Context-Aware Framework for Causal Target Sound Extraction
Shrishail Baligar, Mikolaj Kegler, Bryce Irvin +2
Target Sound Extraction (TSE) focuses on the problem of separating sources of interest, indicated by a user's cue, from the input mixture. Most existing solutions operate in an off…
Latent CLAP Loss for Better Foley Sound Synthesis
Tornike Karchkhadze, Hassan Salami Kavaki, Mohammad Rasool Izadi +5
Foley sound generation, the art of creating audio for multimedia, has recently seen notable advancements through text-conditioned latent diffusion models. These systems use multimo…
CCATMos: Convolutional Context-aware Transformer Network for Non-intrusive Speech Quality Assessment
Yuchen Liu, Li-Chia Yang, Alex Pawlicki +1
Speech quality assessment has been a critical component in many voice communication related applications such as telephony and online conferencing. Traditional intrusive speech qua…
Self-Supervised Learning for Speech Enhancement through Synthesis
Bryce Irvin, Marko Stamenovic, Mikolaj Kegler +1
Modern speech enhancement (SE) networks typically implement noise suppression through time-frequency masking, latent representation masking, or discriminative signal prediction. In…