activity
20182026
most citedAcoustic Scene Classification Using Fusion of Attentive Convolutional Neural Networks for DCASE2019 Challenge

6 citations · 21 across the 24 of their papers we have counts for

collaborators
Showing 2024Show all

6 papers · 1 filter

eess.AS2024

DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition

Alexander Polok, Dominik Klement, Martin Kocour +7

Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a significant challenge, particularly when systems conditioned on speaker embeddings fai…

eess.AS2024

Target Speaker ASR with Whisper

Alexander Polok, Dominik Klement, Matthew Wiesner +3

We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier t…

eess.AS2024

Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units

Bolaji Yusuf, Jan "Honza" Černocký, Murat Saraçlar

End-to-end (E2E) keyword search (KWS) has emerged as an alternative and complimentary approach to conventional keyword search which depends on the output of automatic speech recogn…

eess.AS2024

Written Term Detection Improves Spoken Term Detection

Bolaji Yusuf, Murat Saraçlar

End-to-end (E2E) approaches to keyword search (KWS) are considerably simpler in terms of training and indexing complexity when compared to approaches which use the output of automa…

eess.AS2024

Probing Self-supervised Learning Models with Target Speech Extraction

Junyi Peng, Marc Delcroix, Tsubasa Ochiai +4

Large-scale pre-trained self-supervised learning (SSL) models have shown remarkable advancements in speech-related tasks. However, the utilization of these models in complex multi-…

eess.AS2024

Target Speech Extraction with Pre-trained Self-supervised Learning Models

Junyi Peng, Marc Delcroix, Tsubasa Ochiai +3

Pre-trained self-supervised learning (SSL) models have achieved remarkable success in various speech tasks. However, their potential in target speech extraction (TSE) has not been…