activity
20242026
collaborators

5 papers

cs.SD2026

Metric Analysis for Spatial Semantic Segmentation of Sound Scenes

Mayank Mishra, Paul Magron, Romain Serizel

Spatial semantic segmentation of sound scenes (S5) consists of jointly performing audio source separation and sound event classification from a multichannel audio mixture. Evaluati…

cs.SD2025

Frequency-Weighted Training Losses for Phoneme-Level DNN-based Speech Enhancement

Nasser-Eddine Monir, Paul Magron, Romain Serizel

Recent advances in deep learning have significantly improved multichannel speech enhancement algorithms, yet conventional training loss functions such as the scale-invariant signal…

cs.SD2025

Evaluating Multichannel Speech Enhancement Algorithms at the Phoneme Scale Across Genders

Nasser-Eddine Monir, Paul Magron, Romain Serizel

Multichannel speech enhancement algorithms are essential for improving the intelligibility of speech signals in noisy environments. These algorithms are usually evaluated at the ut…

cs.SD2025

Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes

Masahiro Yasuda, Binh Thien Nguyen, Noboru Harada +10

Spatial Semantic Segmentation of Sound Scenes (S5) aims to enhance technologies for sound event detection and separation from multi-channel input signals that mix multiple sound ev…

cs.SD2024

Angular Distance Distribution Loss for Audio Classification

Antonio Almudévar, Romain Serizel, Alfonso Ortega

Classification is a pivotal task in deep learning not only because of its intrinsic importance, but also for providing embeddings with desirable properties in other tasks. To optim…