wav2pos: Sound Source Localization using Masked Autoencoders
arXiv:2408.15771 · doi:10.1109/IPIN62893.2024.10786105
Abstract
We present a novel approach to the 3D sound source localization task for distributed ad-hoc microphone arrays by formulating it as a set-to-set regression problem. By training a multi-modal masked autoencoder model that operates on audio recordings and microphone coordinates, we show that such a formulation allows for accurate localization of the sound source, by reconstructing coordinates masked in the input. Our approach is flexible in the sense that a single model can be used with an arbitrary number of microphones, even when a subset of audio recordings and microphone coordinates are missing. We test our method on simulated and real-world recordings of music and speech in indoor environments, and demonstrate competitive performance compared to both classical and other learning based localization methods.
IPIN 2024
References in corpus (5)
- Pyroomacoustics: A Python package for audio room simulations and array processing algorithms
- A Survey of Sound Source Localization with Deep Learning Methods
- Towards End-to-End Acoustic Localization using Deep Learning: from Audio Signal to Source Position Coordinates
- Robust Sound Source Tracking Using SRP-PHAT and 3D Convolutional Neural Networks
- Extending GCC-PHAT using Shift Equivariant Neural Networks