activity
20242026
collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2026

Anomalous Sound Detection Meets Noise-Aware Self-Supervised Learning

Takuya Fujimura, Gordon Wichern, Yoshiki Masuyama +5

In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio re…

eess.AS2026

Technical Report for MERL's Real-TSE Challenge Submission

Dominik Klement, Yoshiki Masuyama, Christoph Boeddeker +4

Target speech extraction (TSE) has largely been dominated by neural network-based approaches trained and evaluated on synthetic fully overlapped data. The Real-TSE Challenge aims t…

eess.AS2026

Input-Adaptive Spectral Feature Compression by Sequence Modeling for Source Separation

Kohei Saijo, Yoshiaki Bando

Time-frequency domain dual-path models have demonstrated strong performance and are widely used in source separation. Because their computational cost grows with the number of freq…

eess.AS2024

Task-Aware Unified Source Separation

Kohei Saijo, Janek Ebbers, François G. Germain +2

Several attempts have been made to handle multiple source separation tasks such as speech enhancement, speech separation, sound event separation, music source separation (MSS), or…

eess.AS2024

Leveraging Audio-Only Data for Text-Queried Target Sound Extraction

Kohei Saijo, Janek Ebbers, François G. Germain +3

The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access…

eess.AS2024

TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement

Kohei Saijo, Gordon Wichern, François G. Germain +2

Time-frequency (TF) domain dual-path models achieve high-fidelity speech separation. While some previous state-of-the-art (SoTA) models rely on RNNs, this reliance means they lack…