activity
20242026
collaborators
Showing eess.ASShow all

15 papers · 1 filter

eess.AS2026

SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization

Petr Pálka, Jiangyu Han, Prachi Singh +3

We propose SphereVBx, a Bayesian clustering framework for hyperspherical embeddings based on Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA). The method follows the…

eess.AS2026

Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning

Alexander Polok, Samuele Cornell, Sathvik Udupa +3

We propose diarization-conditioned spoken language models (SLMs), a strategy for extending SLMs to far-field multi-talker audio. Rather than adapting the decoder via Serialized Out…

eess.AS2026

BUT System Description for CHiME-9 MCoRec Challenge

Dominik Klement, Alexander Polok, Nguyen Hai Phong +2

Multi-talker automatic speech recognition (ASR) in conversational recordings remains an open problem, particularly in scenarios with large portion of overlapping speech where ident…

eess.AS2026

On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization

Marc Deegen, Tobias Gburrek, Tobias Cord-Landwehr +4

Recent advances in speaker diarization exploit large pretrained foundation models, such as WavLM, to achieve state-of-the-art performance on multiple datasets. Systems like DiariZe…

eess.AS2025

State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data

Sara Barahona, Ladislav Mošner, Themos Stafylakis +4

In this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source Vox…

eess.AS2025

Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models

Jiangyu Han, Petr Pálka, Marc Delcroix +4

Self-supervised learning (SSL) models such as WavLM have substantially advanced speaker diarization by providing rich contextual speech representations. However, the high computati…