9 papers · 1 filter
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
Martin Kocour, Martin Karafiat, Alexander Polok +3
We propose a speaker-attributed (SA) Whisper-based model for multi-talker speech recognition that combines target-speaker modeling with serialized output training (SOT). Our approa…
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
Alexander Polok, Dominik Klement, Samuele Cornell +4
Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a major challenge. While some approaches achieve strong performance when fine-tuned on s…
BUT Systems for Environmental Sound Deepfake Detection in the ESDD 2026 Challenge
Junyi Peng, Lin Zhang, Jin Li +2
This paper describes the BUT submission to the ESDD 2026 Challenge, specifically focusing on Track 1: Environmental Sound Deepfake Detection with Unseen Generators. To address the…
Unsupervised Speech Enhancement using Data-defined Priors
Dominik Klement, Matthew Maciejewski, Sanjeev Khudanpur +2
The majority of deep learning-based speech enhancement methods require paired clean-noisy speech data. Collecting such data at scale in real-world conditions is infeasible, which h…
Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
Jiangyu Han, Federico Landini, Johan Rohdin +4
Self-supervised learning (SSL) models like WavLM can be effectively utilized when building speaker diarization systems but are often large and slow, limiting their use in resource…
Target Speaker ASR with Whisper
Alexander Polok, Dominik Klement, Matthew Wiesner +3
We propose a novel approach to enable the use of large, single-speaker ASR models, such as Whisper, for target speaker ASR. The key claim of this method is that it is much easier t…