4 papers · 1 filter
Speaker head orientation estimation with a single microphone array using phase spectrogram features
Balint Turi, Archontis Politis, Parthasaarathy Sudarsanam +1
Estimating a speaker's head orientation from audio can provide valuable information in smart environments, meetings, and driver monitoring. We propose a novel approach that leverag…
Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification
Kazuki Shimada, Archontis Politis, Iran R. Roman +10
This paper presents the objective, dataset, baseline, and metrics of Task 3 of the DCASE2025 Challenge on sound event localization and detection (SELD). In previous editions, the c…
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
Parthasaarathy Sudarsanam, Irene MartÃn-Morató, Tuomas Virtanen
This paper proposes a single-stage training approach that semantically aligns three modalities - audio, visual, and text using a contrastive learning framework. Contrastive trainin…
Knowledge Distillation for Speech Denoising by Latent Representation Alignment with Cosine Distance
Diep Luong, Mikko Heikkinen, Konstantinos Drossos +1
Speech denoising is a generally adopted and impactful task, appearing in many common and everyday-life use cases. Although there are very powerful methods published, most of those…