9 papers
Exploring Efficient Waveform Diffusion Models for Foley Sound Generation
Runwu Shi, Chang Li, Jiahui Li +7
Recent advances in diffusion models have enabled high-fidelity Foley sound generation directly in the waveform space. Existing waveform diffusion models primarily rely on time-doma…
What Do Neural Networks Learn for TDOA Estimation? A Cross-Architecture Probing Study
Yaozhong Kang, Jiang Wang, Runwu Shi +3
Neural networks outperform classical GCC-PHAT for Time-Difference-of-Arrival (TDOA) estimation in noise and reverberation, yet their internal strategy remains unexplored. To uncove…
Fast-SDE: Efficient Single-Microphone Sound Source Distance Estimation in Reverberant Environments
Jiang Wang, Runwu Shi, Yaozhong Kang +3
Sound source distance estimation (SDE) is a critical capability in human-robot interaction. An inappropriate interaction distance not only reduces the reliability of speech acquisi…
ASAP: An Azimuth-Priority Strip-Based Search Approach to Planar Microphone Array DOA Estimation in 3D
Ming Huang, Shuting Xu, Leying Yang +6
Direction-of-arrival (DOA) estimation is an important task in microphone array processing and many downstream applications. The steered response power with phase transform (SRP-PHA…
Unsupervised Single-Channel Speech Separation with Diffusion under Speaker-Embedding Guidance
Runwu Shi, Kai Li, Chang Li +5
Speech separation is a fundamental task in audio processing, typically addressed with fully supervised systems trained on paired mixtures. While effective, such systems typically r…
Single-Microphone-Based Sound Source Localization for Mobile Robots in Reverberant Environments
Jiang Wang, Runwu Shi, Benjamin Yen +2
Accurately estimating sound source positions is crucial for robot audition. However, existing sound source localization methods typically rely on a microphone array with at least t…