10 papers · 1 filter
On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization
Marc Deegen, Tobias Gburrek, Tobias Cord-Landwehr +4
Recent advances in speaker diarization exploit large pretrained foundation models, such as WavLM, to achieve state-of-the-art performance on multiple datasets. Systems like DiariZe…
VBx for End-to-End Neural and Clustering-based Diarization
Petr Pálka, Jiangyu Han, Marc Delcroix +2
We present improvements to speaker diarization in the two-stage end-to-end neural diarization with vector clustering (EEND-VC) framework. The first stage employs a Conformer-based…
Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization
Jiangyu Han, Ruoyu Wang, Yoshiki Masuyama +4
Self-supervised models such as WavLM have demonstrated strong performance for neural speaker diarization. However, these models are typically pre-trained on single-channel recordin…
Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing
Junyi Peng, Lin Zhang, Jiangyu Han +5
Although large-scale self-supervised learning (SSL) models like WavLM have achieved state-of-the-art performance in speech processing, their significant size impedes deployment on…
BUT System for the MLC-SLM Challenge
Alexander Polok, Jiangyu Han, Dominik Klement +3
We present a two-speaker automatic speech recognition (ASR) system that combines DiCoW -- a diarization-conditioned variant of Whisper -- with DiariZen, a diarization pipeline buil…
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
Jiangyu Han, Petr Pálka, Marc Delcroix +4
Self-supervised learning (SSL) models such as WavLM have substantially advanced speaker diarization by providing rich contextual speech representations. However, the high computati…