17 papers
Dynamic Clustering for Cross-Segment Permutation Alignment in Long Speech Separation
Yuzhu Wang, Archontis Politis, Konstantinos Drossos +1
Long speech separation typically employs a segment-separation-stitch paradigm where recordings are divided into short segments, processed independently, and stitched together. Its…
Neural Array-Generic Direction-of-Arrival Estimation Exploiting Array Transfer Functions
Mikko Heikkinen, Archontis Politis, Konstantinos Drossos +1
Direction-of-arrival (DoA) estimation is a key component of multichannel audio processing, yet many deep learning approaches remain tied to the microphone arrays used during traini…
Speaker head orientation estimation with a single microphone array using phase spectrogram features
Balint Turi, Archontis Politis, Parthasaarathy Sudarsanam +1
Estimating a speaker's head orientation from audio can provide valuable information in smart environments, meetings, and driver monitoring. We propose a novel approach that leverag…
CNN Models for Microphone Array Covariance Matrix Upsampling and Acoustic Imaging
Marianthi Adamopoulou, Parthasaarathy Sudarsanam, David Diaz-Guerra +5
Acoustic imaging visualization is a core methodology in acoustics, enabling spatial analysis of sound sources and acoustic scenes. However, limited sensor availability in practical…
Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation
Michael Neri, Archontis Politis, Tuomas Virtanen
Single-channel speaker distance estimation has recently achieved centimeter-level accuracy in simulated environments, yet it remains unclear which components of the room impulse re…
Generalised Transcoding Framework for Arbitrary Spatial Audio Capture and Playback Formats
Archontis Politis, Janani Fernandez, Leo McCormack
This article introduces a unified framework for the parametric analysis and reproduction of spatial sound scenes captured either as Ambisonic signals or as raw microphone array sig…