14 papers
Generating Training Targets for Real-World Speech Enhancement via Close-to-Distant Microphone Projection
Tomohiro Nakatani, Rintaro Ikeshita, Naoyuki Kamo +2
Training neural networks (NNs) for speech enhancement (SE) in distant speech-capturing scenarios requires paired distorted and clean reference speech signals. While such data are o…
Description and Discussion on DCASE 2026 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
Binh Thien Nguyen, Masahiro Yasuda, Noboru Harada +8
This paper presents an overview of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2026 Challenge Task 4, Spatial Semantic Segmentation of Sound Scenes (S5).…
Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits
Gilad Nurko, Roi Benita, Yehoshua Dissen +4
Robust classification in noisy environments remains a fundamental challenge in machine learning. Standard approaches typically treat signal enhancement and classification as separa…
Reference Microphone Selection for Guided Source Separation based on the Normalized L-p Norm
Anselm Lohmann, Tomohiro Nakatani, Rintaro Ikeshita +3
Guided Source Separation (GSS) is a popular front-end for distant automatic speech recognition (ASR) systems using spatially distributed microphones. When considering spatially dis…
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
Naoyuki Kamo, Naohiro Tawara, Atsushi Ando +15
In this paper, we introduce a multi-talker distant automatic speech recognition (DASR) system we designed for the DASR task 1 of the CHiME-8 challenge. Our system performs speaker…
Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
Alexis Plaquet, Naohiro Tawara, Marc Delcroix +4
End-to-End Neural Diarization with Vector Clustering is a powerful and practical approach to perform Speaker Diarization. Multiple enhancements have been proposed for the segmentat…