13 papers · 1 filter
Anomalous Sound Detection Meets Noise-Aware Self-Supervised Learning
Takuya Fujimura, Gordon Wichern, Yoshiki Masuyama +5
In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio re…
NABEATs: Noise-Aware Audio Representation Learning
Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern +3
We propose the concept of noise-aware audio self-supervised learning (SSL), whose goal is to encode audio mixtures while suppressing undesired noise, and present Noise-Aware BEATs…
Technical Report for MERL's Real-TSE Challenge Submission
Dominik Klement, Yoshiki Masuyama, Christoph Boeddeker +4
Target speech extraction (TSE) has largely been dominated by neural network-based approaches trained and evaluated on synthetic fully overlapped data. The Real-TSE Challenge aims t…
Predictive-Generative Drift Decomposition for Speech Enhancement and Separation
Julius Richter, Yoshiki Masuyama, Christoph Boeddeker +3
We propose a plug-and-play framework for speech enhancement and separation that augments predictive methods with a generative speech prior. Our approach, termed Stochastic Interpol…
SUNAC: Source-aware Unified Neural Audio Codec
Ryo Aihara, Yoshiki Masuyama, Francesco Paissan +3
Neural audio codecs (NACs) provide compact representations that can be leveraged in many downstream applications, in particular large language models. Yet most NACs encode mixtures…
Handling Domain Shifts for Anomalous Sound Detection: A Review of DCASE-Related Work
Kevin Wilkinghoff, Takuya Fujimura, Keisuke Imoto +3
When detecting anomalous sounds in complex environments, one of the main difficulties is that trained models must be sensitive to subtle differences in monitored target signals, wh…