activity
20192026
most citedJointly optimal denoising, dereverberation, and source separation

45 citations · 134 across the 62 of their papers we have counts for

collaborators
Showing cs.SDShow all

14 papers · 1 filter

cs.SD2026

Frontend Token Enhancement for Token-Based Speech Recognition

Takanori Ashihara, Shota Horiguchi, Kohei Matsuura +2

Discretized representations of speech signals are efficient alternatives to continuous features for various speech applications, including automatic speech recognition (ASR) and sp…

cs.SD2025

Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering

Alexis Plaquet, Naohiro Tawara, Marc Delcroix +4

End-to-End Neural Diarization with Vector Clustering is a powerful and practical approach to perform Speaker Diarization. Multiple enhancements have been proposed for the segmentat…

cs.SD2025

Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes

Masahiro Yasuda, Binh Thien Nguyen, Noboru Harada +10

Spatial Semantic Segmentation of Sound Scenes (S5) aims to enhance technologies for sound event detection and separation from multi-channel input signals that mix multiple sound ev…

cs.SD2024

Investigation of Speaker Representation for Target-Speaker Speech Processing

Takanori Ashihara, Takafumi Moriya, Shota Horiguchi +5

Target-speaker speech processing (TS) tasks, such as target-speaker automatic speech recognition (TS-ASR), target speech extraction (TSE), and personal voice activity detection (p-…

cs.SD2024

Mamba-based Segmentation Model for Speaker Diarization

Alexis Plaquet, Naohiro Tawara, Marc Delcroix +3

Mamba is a newly proposed architecture which behaves like a recurrent neural network (RNN) with attention-like capabilities. These properties are promising for speaker diarization,…

cs.SD2024

SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model

Carlos Hernandez-Olivan, Marc Delcroix, Tsubasa Ochiai +4

Target sound extraction (TSE) consists of isolating a desired sound from a mixture of arbitrary sounds using clues to identify it. A TSE system requires solving two problems at onc…