5 papers · 1 filter
Exploring Efficient Waveform Diffusion Models for Foley Sound Generation
Runwu Shi, Chang Li, Jiahui Li +7
Recent advances in diffusion models have enabled high-fidelity Foley sound generation directly in the waveform space. Existing waveform diffusion models primarily rely on time-doma…
ASAP: An Azimuth-Priority Strip-Based Search Approach to Planar Microphone Array DOA Estimation in 3D
Ming Huang, Shuting Xu, Leying Yang +6
Direction-of-arrival (DOA) estimation is an important task in microphone array processing and many downstream applications. The steered response power with phase transform (SRP-PHA…
Unsupervised Single-Channel Speech Separation with Diffusion under Speaker-Embedding Guidance
Runwu Shi, Kai Li, Chang Li +5
Speech separation is a fundamental task in audio processing, typically addressed with fully supervised systems trained on paired mixtures. While effective, such systems typically r…
Single-Channel Target Speech Extraction Utilizing Distance and Room Clues
Runwu Shi, Zirui Lin, Benjamin Yen +3
This paper aims to achieve single-channel target speech extraction (TSE) in enclosures utilizing distance clues and room information. Recent works have verified the feasibility of…
Asynchronous Microphone Array Calibration using Hybrid TDOA Information
Chengjie Zhang, Jiang Wang, He Kong
Asynchronous microphone array calibration is a prerequisite for many audition robot applications. A popular solution to the above calibration problem is the batch form of Simultane…