activity
20232026
most citedScenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction

1 citations · 1 across the 16 of their papers we have counts for

collaborators

19 papers

eess.AS2026

Downstream-Task-Aware Unified Source Separation

Yoshiki Mitsui, Ryo Aihara, Tatsuhiko Saito +5

Task-aware unified source separation (TUSS) enables a single model to handle diverse separation tasks by conditioning on input prompts. However, conventional TUSS does not account…

eess.AS2026

Anomalous Sound Detection Meets Noise-Aware Self-Supervised Learning

Takuya Fujimura, Gordon Wichern, Yoshiki Masuyama +5

In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio re…

eess.AS2026

NABEATs: Noise-Aware Audio Representation Learning

Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern +3

We propose the concept of noise-aware audio self-supervised learning (SSL), whose goal is to encode audio mixtures while suppressing undesired noise, and present Noise-Aware BEATs…

eess.AS2026

Technical Report for MERL's Real-TSE Challenge Submission

Dominik Klement, Yoshiki Masuyama, Christoph Boeddeker +4

Target speech extraction (TSE) has largely been dominated by neural network-based approaches trained and evaluated on synthetic fully overlapped data. The Real-TSE Challenge aims t…

eess.AS2026

Predictive-Generative Drift Decomposition for Speech Enhancement and Separation

Julius Richter, Yoshiki Masuyama, Christoph Boeddeker +3

We propose a plug-and-play framework for speech enhancement and separation that augments predictive methods with a generative speech prior. Our approach, termed Stochastic Interpol…

cs.SD2026

Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling

Yoshiki Masuyama, Francois G. Germain, Gordon Wichern +2

First-order Ambisonics (FOA) is a standard spatial audio format based on spherical harmonic decomposition. Its zeroth- and first-order components capture the sound pressure and par…