6 citations · 21 across the 24 of their papers we have counts for
6 papers · 1 filter
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
Alexander Polok, Dominik Klement, Martin Kocour +7
Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a significant challenge, particularly when systems conditioned on speaker embeddings fai…
Target Speaker ASR with Whisper
Alexander Polok, Dominik Klement, Matthew Wiesner +3
We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier t…
Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units
Bolaji Yusuf, Jan "Honza" Černocký, Murat Saraçlar
End-to-end (E2E) keyword search (KWS) has emerged as an alternative and complimentary approach to conventional keyword search which depends on the output of automatic speech recogn…
Written Term Detection Improves Spoken Term Detection
Bolaji Yusuf, Murat Saraçlar
End-to-end (E2E) approaches to keyword search (KWS) are considerably simpler in terms of training and indexing complexity when compared to approaches which use the output of automa…
Probing Self-supervised Learning Models with Target Speech Extraction
Junyi Peng, Marc Delcroix, Tsubasa Ochiai +4
Large-scale pre-trained self-supervised learning (SSL) models have shown remarkable advancements in speech-related tasks. However, the utilization of these models in complex multi-…
Target Speech Extraction with Pre-trained Self-supervised Learning Models
Junyi Peng, Marc Delcroix, Tsubasa Ochiai +3
Pre-trained self-supervised learning (SSL) models have achieved remarkable success in various speech tasks. However, their potential in target speech extraction (TSE) has not been…