5 papers · 1 filter
Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
Yoshiki Masuyama, Francois G. Germain, Gordon Wichern +2
First-order Ambisonics (FOA) is a standard spatial audio format based on spherical harmonic decomposition. Its zeroth- and first-order components capture the sound pressure and par…
FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
Yoshiki Masuyama, Kohei Saijo, Francesco Paissan +6
Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configur…
FasTUSS: Faster Task-Aware Unified Source Separation
Francesco Paissan, Gordon Wichern, Yoshiki Masuyama +4
Time-Frequency (TF) dual-path models are currently among the best performing audio source separation network architectures, achieving state-of-the-art performance in speech enhance…
Physics-Informed Direction-Aware Neural Acoustic Fields
Yoshiki Masuyama, François G. Germain, Gordon Wichern +2
This paper presents a physics-informed neural network (PINN) for modeling first-order Ambisonic (FOA) room impulse responses (RIRs). PINNs have demonstrated promising performance i…
SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers
Junghyun Koo, Gordon Wichern, Francois G. Germain +2
We introduce Self-Monitored Inference-Time INtervention (SMITIN), an approach for controlling an autoregressive generative music transformer using classifier probes. These simple l…