12 papers
Anomalous Sound Detection Meets Noise-Aware Self-Supervised Learning
Takuya Fujimura, Gordon Wichern, Yoshiki Masuyama +5
In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio re…
NABEATs: Noise-Aware Audio Representation Learning
Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern +3
We propose the concept of noise-aware audio self-supervised learning (SSL), whose goal is to encode audio mixtures while suppressing undesired noise, and present Noise-Aware BEATs…
Technical Report for MERL's Real-TSE Challenge Submission
Dominik Klement, Yoshiki Masuyama, Christoph Boeddeker +4
Target speech extraction (TSE) has largely been dominated by neural network-based approaches trained and evaluated on synthetic fully overlapped data. The Real-TSE Challenge aims t…
Predictive-Generative Drift Decomposition for Speech Enhancement and Separation
Julius Richter, Yoshiki Masuyama, Christoph Boeddeker +3
We propose a plug-and-play framework for speech enhancement and separation that augments predictive methods with a generative speech prior. Our approach, termed Stochastic Interpol…
Do We Need EMA for Diffusion-Based Speech Enhancement? Toward a Magnitude-Preserving Network Architecture
Julius Richter, Danilo de Oliveira, Timo Gerkmann
We study diffusion-based speech enhancement using a Schrodinger bridge formulation and extend the EDM2 framework to this setting. We employ time-dependent preconditioning of networ…
LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models
Julius Richter, Danilo de Oliveira, Tal Peer +1
We present LipDiffuser, a conditional diffusion model for lip-to-speech generation synthesizing natural and intelligible speech directly from silent video recordings. Our approach…