activity
20182026
most citedLeveraging Low-Distortion Target Estimates for Improved Speech Enhancement

12 citations · 20 across the 28 of their papers we have counts for

collaborators
Showing cs.SDShow all

17 papers · 1 filter

cs.SD2025

FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement

Yoshiki Masuyama, Kohei Saijo, Francesco Paissan +6

Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configur…

cs.SD2025

FasTUSS: Faster Task-Aware Unified Source Separation

Francesco Paissan, Gordon Wichern, Yoshiki Masuyama +4

Time-Frequency (TF) dual-path models are currently among the best performing audio source separation network architectures, achieving state-of-the-art performance in speech enhance…

cs.SD2025

Physics-Informed Direction-Aware Neural Acoustic Fields

Yoshiki Masuyama, François G. Germain, Gordon Wichern +2

This paper presents a physics-informed neural network (PINN) for modeling first-order Ambisonic (FOA) room impulse responses (RIRs). PINNs have demonstrated promising performance i…

cs.SD2024

SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers

Junghyun Koo, Gordon Wichern, Francois G. Germain +2

We introduce Self-Monitored Inference-Time INtervention (SMITIN), an approach for controlling an autoregressive generative music transformer using classifier probes. These simple l…

cs.SD2023

Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation

Shih-Lun Wu, Xuankai Chang, Gordon Wichern +4

Automated audio captioning (AAC) aims to generate informative descriptions for various sounds from nature and/or human activities. In recent years, AAC has quickly attracted resear…

cs.SD2022

Optimal Condition Training for Target Source Separation

Efthymios Tzinis, Gordon Wichern, Paris Smaragdis +1

Recent research has shown remarkable performance in leveraging multiple extraneous conditional and non-mutually exclusive semantic concepts for sound source separation, allowing th…