most citedNeural Target Speech Extraction: An Overview

120 citations · 144 across the 6 of their papers we have counts for

collaborators

11 papers

cs.SD20241 cited

Interaural time difference loss for binaural target sound extraction

Carlos Hernandez-Olivan, Marc Delcroix, Tsubasa Ochiai +3

Binaural target sound extraction (TSE) aims to extract a desired sound from a binaural mixture of arbitrary sounds while preserving the spatial cues of the desired sound. Indeed, f…

eess.AS2024

SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling

Hiroshi Sato, Takafumi Moriya, Masato Mimura +6

Real-time target speaker extraction (TSE) is intended to extract the desired speaker's voice from the observed mixture of multiple speakers in a streaming manner. Implementing real…

eess.AS2024

Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance

Tsubasa Ochiai, Kazuma Iwamoto, Marc Delcroix +4

It is challenging to improve automatic speech recognition (ASR) performance in noisy conditions with a single-channel speech enhancement (SE) front-end. This is generally attribute…

eess.AS2024

Probing Self-supervised Learning Models with Target Speech Extraction

Junyi Peng, Marc Delcroix, Tsubasa Ochiai +4

Large-scale pre-trained self-supervised learning (SSL) models have shown remarkable advancements in speech-related tasks. However, the utilization of these models in complex multi-…

eess.AS2024

Target Speech Extraction with Pre-trained Self-supervised Learning Models

Junyi Peng, Marc Delcroix, Tsubasa Ochiai +3

Pre-trained self-supervised learning (SSL) models have achieved remarkable success in various speech tasks. However, their potential in target speech extraction (TSE) has not been…

eess.AS20231 cited

Knowledge Distillation for Neural Transducer-based Target-Speaker ASR: Exploiting Parallel Mixture/Single-Talker Speech Data

Takafumi Moriya, Hiroshi Sato, Tsubasa Ochiai +7

Neural transducer (RNNT)-based target-speaker speech recognition (TS-RNNT) directly transcribes a target speaker's voice from a multi-talker mixture. It is a promising approach for…