6 papers
Speaker head orientation estimation with a single microphone array using phase spectrogram features
Balint Turi, Archontis Politis, Parthasaarathy Sudarsanam +1
Estimating a speaker's head orientation from audio can provide valuable information in smart environments, meetings, and driver monitoring. We propose a novel approach that leverag…
CNN Models for Microphone Array Covariance Matrix Upsampling and Acoustic Imaging
Marianthi Adamopoulou, Parthasaarathy Sudarsanam, David Diaz-Guerra +5
Acoustic imaging visualization is a core methodology in acoustics, enabling spatial analysis of sound sources and acoustic scenes. However, limited sensor availability in practical…
FOA Tokenizer: Low-bitrate Neural Codec for First Order Ambisonics with Spatial Consistency Loss
Parthasaarathy Sudarsanam, Sebastian Braun, Hannes Gamper
Neural audio codecs have been widely studied for mono and stereo signals, but spatial audio remains largely unexplored. We present the first discrete neural spatial audio codec for…
Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification
Kazuki Shimada, Archontis Politis, Iran R. Roman +10
This paper presents the objective, dataset, baseline, and metrics of Task 3 of the DCASE2025 Challenge on sound event localization and detection (SELD). In previous editions, the c…
Towards Spatial Audio Understanding via Question Answering
Parthasaarathy Sudarsanam, Archontis Politis
In this paper, we introduce a novel framework for spatial audio understanding of first-order ambisonic (FOA) signals through a question answering (QA) paradigm, aiming to extend th…
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
Parthasaarathy Sudarsanam, Irene MartÃn-Morató, Tuomas Virtanen
This paper proposes a single-stage training approach that semantically aligns three modalities - audio, visual, and text using a contrastive learning framework. Contrastive trainin…