activity
20202026
most citedCCATMos: Convolutional Context-aware Transformer Network for Non-intrusive Speech Quality Assessment

4 citations · 12 across the 8 of their papers we have counts for

collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2026

FSD50K-Solo: Automated Curation of Single-Source Sound Events

Ningyuan Yang, Sile Yin, Li-Chia Yang +4

High-quality training datasets are essential for the performance of neural networks. However, the audio domain still lacks a large-scale, strongly-labeled, and single-source sound…

eess.AS2025

PAS-SE: Personalized Auxiliary-Sensor Speech Enhancement for Voice Pickup in Hearables

Mattes Ohlenbusch, Mikolaj Kegler, Marko Stamenovic

Speech enhancement for voice pickup in hearables aims to improve the user's voice by suppressing noise and interfering talkers, while maintaining own-voice quality. For single-chan…

eess.AS2024

CATSE: A Context-Aware Framework for Causal Target Sound Extraction

Shrishail Baligar, Mikolaj Kegler, Bryce Irvin +2

Target Sound Extraction (TSE) focuses on the problem of separating sources of interest, indicated by a user's cue, from the input mixture. Most existing solutions operate in an off…

eess.AS2024

Latent CLAP Loss for Better Foley Sound Synthesis

Tornike Karchkhadze, Hassan Salami Kavaki, Mohammad Rasool Izadi +5

Foley sound generation, the art of creating audio for multimedia, has recently seen notable advancements through text-conditioned latent diffusion models. These systems use multimo…

eess.AS20224 cited

CCATMos: Convolutional Context-aware Transformer Network for Non-intrusive Speech Quality Assessment

Yuchen Liu, Li-Chia Yang, Alex Pawlicki +1

Speech quality assessment has been a critical component in many voice communication related applications such as telephony and online conferencing. Traditional intrusive speech qua…

eess.AS20223 cited

Self-Supervised Learning for Speech Enhancement through Synthesis

Bryce Irvin, Marko Stamenovic, Mikolaj Kegler +1

Modern speech enhancement (SE) networks typically implement noise suppression through time-frequency masking, latent representation masking, or discriminative signal prediction. In…