Pyroomacoustics: A Python package for audio room simulations and array processing algorithms
arXiv:1710.04196 · doi:10.1109/ICASSP.2018.8461310
Abstract
We present pyroomacoustics, a software package aimed at the rapid development and testing of audio array processing algorithms. The content of the package can be divided into three main components: an intuitive Python object-oriented interface to quickly construct different simulation scenarios involving multiple sound sources and microphones in 2D and 3D rooms; a fast C implementation of the image source model for general polyhedral rooms to efficiently generate room impulse responses and simulate the propagation between sources and receivers; and finally, reference implementations of popular algorithms for beamforming, direction finding, and adaptive filtering. Together, they form a package with the potential to speed up the time to market of new algorithms by significantly reducing the implementation overhead in the performance evaluation step.
5 pages, 5 figures, describes a software package
Cited by in corpus (80)
- A Survey of Sound Source Localization with Deep Learning Methods
- Speech Enhancement and Dereverberation with Diffusion-based Generative Models
- gpuRIR: A Python Library for Room Impulse Response Simulation with GPU Acceleration
- Building and Evaluation of a Real Room Impulse Response Dataset
- StoRM: A Diffusion-based Stochastic Regeneration Model for Speech Enhancement and Dereverberation
- Insights Into Deep Non-linear Filters for Improved Multi-channel Speech Enhancement
- Hearing What You Cannot See: Acoustic Vehicle Detection Around Corners
- ClearBuds: Wireless Binaural Earbuds for Learning-Based Speech Enhancement
- DNN-based mask estimation for distributed speech enhancement in spatially unconstrained microphone arrays
- Semi-supervised source localization with deep generative modeling
- Independent Vector Analysis via Log-Quadratically Penalized Quadratic Minimization
- Multi-channel Speech Separation Using Spatially Selective Deep Non-linear Filters
- BeamLearning: an end-to-end Deep Learning approach for the angular localization of sound sources using raw multichannel acoustic pressure data
- Joint Dereverberation and Separation with Iterative Source Steering
- Ensemble of ACCDOA- and EINV2-based Systems with D3Nets and Impulse Response Simulation for Sound Event Localization and Detection
- Perceptual Based Adversarial Audio Attacks
- Extending GCC-PHAT using Shift Equivariant Neural Networks
- Signal-Aware Direction-of-Arrival Estimation Using Attention Mechanisms
- Localizing Unsynchronized Sensors with Unknown Sources
- Separake: Source Separation with a Little Help From Echoes
- DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing
- Mean absorption estimation from room impulse responses using virtually supervised learning
- Real-time, Universal, and Robust Adversarial Attacks Against Speaker Recognition Systems
- MM Algorithms for Joint Independent Subspace Analysis with Application to Blind Single and Multi-Source Extraction
- Real-time Denoising and Dereverberation with Tiny Recurrent U-Net
- Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning Mechanism
- A Review on Sound Source Localization in Robotics: Focusing on Deep Learning Methods
- Multi-modal Blind Source Separation with Microphones and Blinkies
- Small-Footprint Open-Vocabulary Keyword Spotting with Quantized LSTM Networks
- EchoScan: Scanning Complex Room Geometries via Acoustic Echoes
- Phase-aware Single-stage Speech Denoising and Dereverberation with U-Net
- A Strongly-Labelled Polyphonic Dataset of Urban Sounds with Spatiotemporal Context
- Deep Audio Waveform Prior
- BERP: A Blind Estimator of Room Parameters for Single-Channel Noisy Speech Signals
- SSLIDE: Sound Source Localization for Indoors based on Deep Learning
- wav2pos: Sound Source Localization using Masked Autoencoders
- Distributed speech separation in spatially unconstrained microphone arrays
- Towards Robust Waveform-Based Acoustic Models
- W-Net BF: DNN-based Beamformer Using Joint Training Approach
- Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localization
- Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
- Is Audio Spoof Detection Robust to Laundering Attacks?
- Fully Reversing the Shoebox Image Source Method: From Impulse Responses to Room Parameters
- RGI-Net: 3D Room Geometry Inference from Room Impulse Responses With Hidden First-Order Reflections
- dEchorate: a Calibrated Room Impulse Response Database for Echo-aware Signal Processing
- MIMO-DoAnet: Multi-channel Input and Multiple Outputs DoA Network with Unknown Number of Sound Sources
- CleanUMamba: A Compact Mamba Network for Speech Denoising using Channel Pruning
- Refinement of Direction of Arrival Estimators by Majorization-Minimization Optimization on the Array Manifold
- The INTERSPEECH 2020 Far-Field Speaker Verification Challenge
- Independent Vector Analysis with more Microphones than Sources
- Data-Efficient Framework for Real-world Multiple Sound Source 2D Localization
- Convolutive Prediction for Monaural Speech Dereverberation and Noisy-Reverberant Speaker Separation
- Accoustate: Auto-annotation of IMU-generated Activity Signatures under Smart Infrastructure
- Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments
- Structural sparsification for Far-field Speaker Recognition with GNA
- A study on more realistic room simulation for far-field keyword spotting
- Enhancement by postfiltering for speech and audio coding in ad-hoc sensor networks
- Synthesis of Soundfields through Irregular Loudspeaker Arrays Based on Convolutional Neural Networks
- Room-acoustic simulations as an alternative to measurements for audio-algorithm evaluation
- Fast Independent Vector Extraction by Iterative SINR Maximization
- The DKU System for the Speaker Recognition Task of the 2019 VOiCES from a Distance Challenge
- Short-time deep-learning based source separation for speech enhancement in reverberant environments with beamforming
- HI-MIA : A Far-field Text-Dependent Speaker Verification Database and the Baselines
- Ensemble of Discriminators for Domain Adaptation in Multiple Sound Source 2D Localization
- CLC: Complex Linear Coding for the DNS 2020 Challenge
- OtoWorld: Towards Learning to Separate by Learning to Move
- Cortical Features for Defense Against Adversarial Audio Attacks
- Exploiting Single-Channel Speech For Multi-channel End-to-end Speech Recognition
- RIR-Former: Coordinate-Guided Transformer for Continuous Reconstruction of Room Impulse Responses
- Gridless 3D Recovery of Image Sources from Room Impulse Responses
- BERT for Joint Multichannel Speech Dereverberation with Spatial-aware Tasks
- Speech-dependent Data Augmentation for Own Voice Reconstruction with Hearable Microphones in Noisy Environments
- Resource-Efficient Speech Mask Estimation for Multi-Channel Speech Enhancement
- Length- and Noise-aware Training Techniques for Short-utterance Speaker Recognition
- How much to Dereverberate? Low-Latency Single-Channel Speech Enhancement in Distant Microphone Scenarios
- Independent vector analysis -- an introduction for statisticians
- Surrogate Source Model Learning for Determined Source Separation
- Far-Field Automatic Speech Recognition
- AILive Mixer: A Deep Learning based Zero Latency Automatic Music Mixer for Live Music Performances
- Online Self-Attentive Gated RNNs for Real-Time Speaker Separation