papers

Publications (63)

cs.SD2025

FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement

Yoshiki Masuyama, Kohei Saijo, Francesco Paissan +6

Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configur…

eess.AS2025

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement

Amir Hussein, Sameer Khurana, Gordon Wichern +2

Effective speech representations for spoken language models must balance semantic relevance with acoustic fidelity for high-quality reconstruction. However, existing approaches str…

cs.SD2020

Finding Strength in Weakness: Learning to Separate Sounds with Weak Supervision

Fatemeh Pishdadian, Gordon Wichern, Jonathan Le Roux

While there has been much recent progress using deep learning techniques to separate speech and music audio signals, these systems typically require large collections of isolated s…

eess.AS2026

NABEATs: Noise-Aware Audio Representation Learning

Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern +3

We propose the concept of noise-aware audio self-supervised learning (SSL), whose goal is to encode audio mixtures while suppressing undesired noise, and present Noise-Aware BEATs…

cs.SD2019

Bootstrapping deep music separation from primitive auditory grouping principles

Prem Seetharaman, Gordon Wichern, Jonathan Le Roux +1

Separating an audio scene such as a cocktail party into constituent, meaningful components is a core task in computer audition. Deep networks are the state-of-the-art approach. The…

eess.AS2025

Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations

Ryo Aihara, Yoshiki Masuyama, Gordon Wichern +2

Neural audio codecs (NACs), which use neural networks to generate compact audio representations, have garnered interest for their applicability to many downstream tasks -- especial…

eess.AS2023

NeuroHeed+: Improving Neuro-steered Speaker Extraction with Joint Auditory Attention Detection

Zexu Pan, Gordon Wichern, Francois G. Germain +2

Neuro-steered speaker extraction aims to extract the listener's brain-attended speech signal from a multi-talker speech signal, in which the attention is derived from the cortical…

cs.SD2018

Class-conditional embeddings for music source separation

Prem Seetharaman, Gordon Wichern, Shrikant Venkataramani +1

Isolating individual instruments in a musical mixture has a myriad of potential applications, and seems imminently achievable given the levels of performance reached by recent deep…

eess.AS2025

Local Density-Based Anomaly Score Normalization for Domain Generalization

Kevin Wilkinghoff, Haici Yang, Janek Ebbers +3

State-of-the-art anomalous sound detection (ASD) systems in domain-shifted conditions rely on projecting audio signals into an embedding space and using distance-based outlier dete…

eess.AS2022

Tackling the Cocktail Fork Problem for Separation and Transcription of Real-World Soundtracks

Darius Petermann, Gordon Wichern, Aswin Shanmugam Subramanian +2

Emulating the human ability to solve the cocktail party problem, i.e., focus on a source of interest in a complex acoustic scene, is a long standing goal of audio source separation…

eess.AS2024

Leveraging Audio-Only Data for Text-Queried Target Sound Extraction

Kohei Saijo, Janek Ebbers, François G. Germain +3

The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access…

eess.AS2024

The Sound Demixing Challenge 2023 $\unicode{x2013}$ Cinematic Demixing Track

Stefan Uhlich, Giorgio Fabbro, Masato Hirano +14

This paper summarizes the cinematic demixing (CDX) track of the Sound Demixing Challenge 2023 (SDX'23). We provide a comprehensive summary of the challenge setup, detailing the str…

eess.AS2026

Predictive-Generative Drift Decomposition for Speech Enhancement and Separation

Julius Richter, Yoshiki Masuyama, Christoph Boeddeker +3

We propose a plug-and-play framework for speech enhancement and separation that augments predictive methods with a generative speech prior. Our approach, termed Stochastic Interpol…

eess.AS2024

Task-Aware Unified Source Separation

Kohei Saijo, Janek Ebbers, François G. Germain +2

Several attempts have been made to handle multiple source separation tasks such as speech enhancement, speech separation, sound event separation, music source separation (MSS), or…

cs.SD2021

Leveraging Low-Distortion Target Estimates for Improved Speech Enhancement

Zhong-Qiu Wang, Gordon Wichern, Jonathan Le Roux

A promising approach for multi-microphone speech separation involves two deep neural networks (DNN), where the predicted target speech from the first DNN is used to compute signal…

cs.LG2025

Meta-Learning for Physically-Constrained Neural System Identification

Ankush Chakrabarty, Gordon Wichern, Vedang M. Deshpande +3

We present a gradient-based meta-learning framework for rapid adaptation of neural state-space models (NSSMs) for black-box system identification. When applicable, we also incorpor…

cs.LG2025

Probabilistic Forecasting for Building Energy Systems using Time-Series Foundation Models

Young Jin Park, Francois Germain, Jing Liu +6

Decision-making in building energy systems critically depends on the predictive accuracy of relevant time-series models. In scenarios lacking extensive data from a target building,…

eess.AS2023

Cold Diffusion for Speech Enhancement

Hao Yen, François G. Germain, Gordon Wichern +1

Diffusion models have recently shown promising results for difficult enhancement tasks such as the conditional and unconditional restoration of natural images and audio signals. In…

eess.AS2025

SUNAC: Source-aware Unified Neural Audio Codec

Ryo Aihara, Yoshiki Masuyama, Francesco Paissan +3

Neural audio codecs (NACs) provide compact representations that can be leveraged in many downstream applications, in particular large language models. Yet most NACs encode mixtures…

cs.SD2025

Physics-Informed Direction-Aware Neural Acoustic Fields

Yoshiki Masuyama, François G. Germain, Gordon Wichern +2

This paper presents a physics-informed neural network (PINN) for modeling first-order Ambisonic (FOA) room impulse responses (RIRs). PINNs have demonstrated promising performance i…

eess.AS2020

AutoClip: Adaptive Gradient Clipping for Source Separation Networks

Prem Seetharaman, Gordon Wichern, Bryan Pardo +1

Clipping the gradient is a known approach to improving gradient descent, but requires hand selection of a clipping threshold hyperparameter. We present AutoClip, a simple method fo…

eess.AS2023

Generation or Replication: Auscultating Audio Latent Diffusion Models

Dimitrios Bralios, Gordon Wichern, François G. Germain +4

The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how…

eess.AS2023

Scenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction

Zexu Pan, Gordon Wichern, Yoshiki Masuyama +4

Target speech extraction aims to extract, based on a given conditioning cue, a target speech signal that is corrupted by interfering sources, such as noise or competing speakers. B…

cs.SD2020

Transcription Is All You Need: Learning to Separate Musical Mixtures with Score as Supervision

Yun-Ning Hung, Gordon Wichern, Jonathan Le Roux

Most music source separation systems require large collections of isolated sources for training, which can be difficult to obtain. In this work, we use musical scores, which are co…

eess.AS2024

TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement

Kohei Saijo, Gordon Wichern, François G. Germain +2

Time-frequency (TF) domain dual-path models achieve high-fidelity speech separation. While some previous state-of-the-art (SoTA) models rely on RNNs, this reliance means they lack…

cs.SD2021

On The Compensation Between Magnitude and Phase in Speech Separation

Zhong-Qiu Wang, Gordon Wichern, Jonathan Le Roux

Deep neural network (DNN) based end-to-end optimization in the complex time-frequency (T-F) domain or time domain has shown considerable potential in monaural speech separation. Ma…

cs.SD2019

Phasebook and Friends: Leveraging Discrete Representations for Source Separation

Jonathan Le Roux, Gordon Wichern, Shinji Watanabe +2

Deep learning based speech enhancement and source separation systems have recently reached unprecedented levels of quality, to the point that performance is reaching a new ceiling.…

eess.AS2025

Factorized RVQ-GAN For Disentangled Speech Tokenization

Sameer Khurana, Dominik Klement, Antoine Laurent +13

We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single…

eess.AS2026

Mind the Gap: Detecting Cluster Exits for Robust Local Density-Based Score Normalization in Anomalous Sound Detection

Kevin Wilkinghoff, Gordon Wichern, Jonathan Le Roux +1

Local density-based score normalization is an effective component of distance-based embedding methods for anomalous sound detection, particularly when data densities vary across co…

eess.AS2024

Sound Event Bounding Boxes

Janek Ebbers, Francois G. Germain, Gordon Wichern +1

Sound event detection is the task of recognizing sounds and determining their extent (onset/offset times) within an audio clip. Existing systems commonly predict sound presence con…

cs.LG2021

Attentive Neural Processes and Batch Bayesian Optimization for Scalable Calibration of Physics-Informed Digital Twins

Ankush Chakrabarty, Gordon Wichern, Christopher Laughman

Physics-informed dynamical system models form critical components of digital twins of the built environment. These digital twins enable the design of energy-efficient infrastructur…

cs.SD2021

Convolutive Prediction for Reverberant Speech Separation

Zhong-Qiu Wang, Gordon Wichern, Jonathan Le Roux

We investigate the effectiveness of convolutive prediction, a novel formulation of linear prediction for speech dereverberation, for speaker separation in reverberant conditions. T…

eess.AS2024

Enhanced Reverberation as Supervision for Unsupervised Speech Separation

Kohei Saijo, Gordon Wichern, François G. Germain +2

Reverberation as supervision (RAS) is a framework that allows for training monaural speech separation models from multi-channel mixtures in an unsupervised manner. In RAS, models a…

cs.SD2019

Cutting Music Source Separation Some Slakh: A Dataset to Study the Impact of Training Data Quality and Quantity

Ethan Manilow, Gordon Wichern, Prem Seetharaman +1

Music source separation performance has greatly improved in recent years with the advent of approaches based on deep learning. Such methods typically require large amounts of label…

eess.AS2026

Technical Report for MERL's Real-TSE Challenge Submission

Dominik Klement, Yoshiki Masuyama, Christoph Boeddeker +4

Target speech extraction (TSE) has largely been dominated by neural network-based approaches trained and evaluated on synthetic fully overlapped data. The Real-TSE Challenge aims t…

eess.AS2026

Anomalous Sound Detection Meets Noise-Aware Self-Supervised Learning

Takuya Fujimura, Gordon Wichern, Yoshiki Masuyama +5

In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio re…

cs.LG2022

Meta-Learning of Neural State-Space Models Using Data From Similar Systems

Ankush Chakrabarty, Gordon Wichern, Christopher R. Laughman

Deep neural state-space models (SSMs) provide a powerful tool for modeling dynamical systems solely using operational data. Typically, neural SSMs are trained using data collected…

cs.SD2025

FasTUSS: Faster Task-Aware Unified Source Separation

Francesco Paissan, Gordon Wichern, Yoshiki Masuyama +4

Time-Frequency (TF) dual-path models are currently among the best performing audio source separation network architectures, achieving state-of-the-art performance in speech enhance…

eess.AS2024

TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings

Christoph Boeddeker, Aswin Shanmugam Subramanian, Gordon Wichern +2

Since diarization and source separation of meeting data are closely related tasks, we here propose an approach to perform the two objectives jointly. It builds upon the target-spea…

cs.SD2023

Pac-HuBERT: Self-Supervised Music Source Separation via Primitive Auditory Clustering and Hidden-Unit BERT

Ke Chen, Gordon Wichern, François G. Germain +1

In spite of the progress in music source separation research, the small amount of publicly-available clean source data remains a constant limiting factor for performance. Thus, rec…

eess.AS2022

Locate This, Not That: Class-Conditioned Sound Event DOA Estimation

Olga Slizovskaia, Gordon Wichern, Zhong-Qiu Wang +1

Existing systems for sound event localization and detection (SELD) typically operate by estimating a source location for all classes at every time instant. In this paper, we propos…

cs.SD2022

Heterogeneous Target Speech Separation

Efthymios Tzinis, Gordon Wichern, Aswin Subramanian +2

We introduce a new paradigm for single-channel target source separation where the sources of interest can be distinguished using non-mutually exclusive concepts (e.g., loudness, ge…

eess.AS2025

Retrieval-Augmented Neural Field for HRTF Upsampling and Personalization

Yoshiki Masuyama, Gordon Wichern, François G. Germain +2

Head-related transfer functions (HRTFs) with dense spatial grids are desired for immersive binaural audio generation, but their recording is time-consuming. Although HRTF spatial u…

cs.SD2020

WHAMR!: Noisy and Reverberant Single-Channel Speech Separation

Matthew Maciejewski, Gordon Wichern, Emmett McQuinn +1

While significant advances have been made with respect to the separation of overlapping speech signals, studies have been largely constrained to mixtures of clean, near anechoic sp…

cs.SD2019

WHAM!: Extending Speech Separation to Noisy Environments

Gordon Wichern, Joe Antognini, Michael Flynn +5

Recent progress in separating the speech signals from multiple overlapping speakers using a single audio channel has brought us closer to solving the cocktail party problem. Howeve…

eess.AS2024

NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization

Yoshiki Masuyama, Gordon Wichern, François G. Germain +4

Head-related transfer functions (HRTFs) are important for immersive audio, and their spatial interpolation has been studied to upsample finite measurements. Recently, neural fields…

eess.AS2025

Data Augmentation Using Neural Acoustic Fields With Retrieval-Augmented Pre-training

Christopher Ick, Gordon Wichern, Yoshiki Masuyama +2

This report details MERL's system for room impulse response (RIR) estimation submitted to the Generative Data Augmentation Workshop at ICASSP 2025 for Augmenting RIR Data (Task 1)…

eess.AS2025

Direction-Aware Neural Acoustic Fields for Few-Shot Interpolation of Ambisonic Impulse Responses

Christopher Ick, Gordon Wichern, Yoshiki Masuyama +2

The characteristics of a sound field are intrinsically linked to the geometric and spatial properties of the environment surrounding a sound source and a listener. The physics of s…

cs.SD2023

Latent Iterative Refinement for Modular Source Separation

Dimitrios Bralios, Efthymios Tzinis, Gordon Wichern +2

Traditional source separation approaches train deep neural network models end-to-end with all the data available at once by minimizing the empirical risk on the whole training set.…

cs.SD2022

Optimal Condition Training for Target Source Separation

Efthymios Tzinis, Gordon Wichern, Paris Smaragdis +1

Recent research has shown remarkable performance in leveraging multiple extraneous conditional and non-mutually exclusive semantic concepts for sound source separation, allowing th…

cs.SD2022

STFT-Domain Neural Speech Enhancement with Very Low Algorithmic Latency

Zhong-Qiu Wang, Gordon Wichern, Shinji Watanabe +1

Deep learning based speech enhancement in the short-time Fourier transform (STFT) domain typically uses a large window length such as 32 ms. A larger window can lead to higher freq…

cs.SD2018

Bootstrapping single-channel source separation via unsupervised spatial clustering on stereo mixtures

Prem Seetharaman, Gordon Wichern, Jonathan Le Roux +1

Separating an audio scene into isolated sources is a fundamental problem in computer audition, analogous to image segmentation in visual scene analysis. Source separation systems b…

eess.AS2025

30+ Years of Source Separation Research: Achievements and Future Challenges

Shoko Araki, Nobutaka Ito, Reinhold Haeb-Umbach +3

Source separation (SS) of acoustic signals is a research field that emerged in the mid-1990s and has flourished ever since. On the occasion of ICASSP's 50th anniversary, we review…

eess.AS2023

Late Audio-Visual Fusion for In-The-Wild Speaker Diarization

Zexu Pan, Gordon Wichern, François G. Germain +2

Speaker diarization is well studied for constrained audios but little explored for challenging in-the-wild videos, which have more speakers, shorter utterances, and inconsistent on…

cs.SD2025

SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers

Junghyun Koo, Gordon Wichern, Francois G. Germain +2

We introduce Self-Monitored Inference-Time INtervention (SMITIN), an approach for controlling an autoregressive generative music transformer using classifier probes. These simple l…

eess.AS2022

Hyperbolic Audio Source Separation

Darius Petermann, Gordon Wichern, Aswin Subramanian +1

We introduce a framework for audio source separation using embeddings on a hyperbolic manifold that compactly represent the hierarchical relationship between sound sources and time…

cs.SD2024

Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation

Shih-Lun Wu, Xuankai Chang, Gordon Wichern +4

Automated audio captioning (AAC) aims to generate informative descriptions for various sounds from nature and/or human activities. In recent years, AAC has quickly attracted resear…

eess.AS2024

Why does music source separation benefit from cacophony?

Chang-Bin Jeon, Gordon Wichern, François G. Germain +1

In music source separation, a standard training data augmentation procedure is to create new training samples by randomly combining instrument stems from different songs. These ran…

eess.AS2022

The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks

Darius Petermann, Gordon Wichern, Zhong-Qiu Wang +1

The cocktail party problem aims at isolating any source of interest within a complex acoustic scene, and has long inspired audio source separation research. Recent efforts have mai…

eess.AS2022

Reverberation as Supervision for Speech Separation

Rohith Aralikatti, Christoph Boeddeker, Gordon Wichern +2

This paper proposes reverberation as supervision (RAS), a novel unsupervised loss function for single-channel reverberant speech separation. Prior methods for unsupervised separati…

cs.SD2021

Convolutive Prediction for Monaural Speech Dereverberation and Noisy-Reverberant Speaker Separation

Zhong-Qiu Wang, Gordon Wichern, Jonathan Le Roux

A promising approach for speech dereverberation is based on supervised learning, where a deep neural network (DNN) is trained to predict the direct sound from noisy-reverberant spe…

cs.CL2018

End-to-End Audio Visual Scene-Aware Dialog using Multimodal Attention-Based Video Features

Chiori Hori, Huda Alamri, Jue Wang +10

Dialog systems need to understand dynamic visual scenes in order to have conversations with users about the objects and events around them. Scene-aware dialog systems for real-worl…

cs.SD2026

Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling

Yoshiki Masuyama, Francois G. Germain, Gordon Wichern +2

First-order Ambisonics (FOA) is a standard spatial audio format based on spherical harmonic decomposition. Its zeroth- and first-order components capture the sound pressure and par…