activity
20202023
most citedNeural Target Speech Extraction: An Overview

120 citations · 139 across the 6 of their papers we have counts for

collaborators

8 papers

eess.AS2023★ 120 cited

Neural Target Speech Extraction: An Overview

Katerina Zmolikova, Marc Delcroix, Tsubasa Ochiai +3

Humans can listen to a target speaker even in challenging acoustic conditions that have noise, reverberation, and interfering speakers. This phenomenon is known as the cocktail-par…

eess.AS2022★ 3 cited

Listen only to me! How well can target speech extraction handle false alarms?

Marc Delcroix, Keisuke Kinoshita, Tsubasa Ochiai +3

Target speech extraction (TSE) extracts the speech of a target speaker in a mixture given auxiliary clues characterizing the speaker, such as an enrollment utterance. TSE addresses…

eess.AS2021

Revisiting joint decoding based multi-talker speech recognition with DNN acoustic model

Martin Kocour, Kateřina Žmolíková, Lucas Ondel +5

In typical multi-talker speech recognition systems, a neural network-based acoustic model predicts senone state posteriors for each speaker. These are later used by a single-talker…

eess.AS2021★ 4 cited

Speaker activity driven neural speech extraction

Marc Delcroix, Katerina Zmolikova, Tsubasa Ochiai +2

Target speech extraction, which extracts the speech of a target speaker in a mixture given auxiliary speaker clues, has recently received increased interest. Various clues have bee…

eess.AS2020

Integration of variational autoencoder and spatial clustering for adaptive multi-channel neural speech separation

Katerina Zmolikova, Marc Delcroix, Lukáš Burget +2

In this paper, we propose a method combining variational autoencoder model of speech with a spatial clustering approach for multi-channel speech separation. The advantage of integr…

cs.SD2020

Jointly Trained Transformers models for Spoken Language Translation

Hari Krishna Vydana, Martin Karafi'at, Katerina Zmolikova +2

Conventional spoken language translation (SLT) systems are pipeline based systems, where we have an Automatic Speech Recognition (ASR) system to convert the modality of source from…