13 papers · 1 filter
SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification
Junyi Peng, Oldřich Plchot, Xiao Song +9
Modern speaker verification (SV) systems rely on speaker embeddings that are effective but difficult to interpret or query in natural language. Most existing speech-text corpora ta…
Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning
Alexander Polok, Samuele Cornell, Sathvik Udupa +3
We propose diarization-conditioned spoken language models (SLMs), a strategy for extending SLMs to far-field multi-talker audio. Rather than adapting the decoder via Serialized Out…
DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
Alexander Polok, Santosh Kesiraju, Karel Beneš +3
This paper presents a simple yet effective regularization for the internal language model induced by the decoder in encoder-decoder ASR models, thereby improving robustness and gen…
Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing
Junyi Peng, Lin Zhang, Jiangyu Han +5
Although large-scale self-supervised learning (SSL) models like WavLM have achieved state-of-the-art performance in speech processing, their significant size impedes deployment on…
BUT System for the MLC-SLM Challenge
Alexander Polok, Jiangyu Han, Dominik Klement +3
We present a two-speaker automatic speech recognition (ASR) system that combines DiCoW -- a diarization-conditioned variant of Whisper -- with DiariZen, a diarization pipeline buil…
Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
Šimon Sedláček, Bolaji Yusuf, Ján Švec +4
In this work, we approach spoken Dialogue State Tracking (DST) by bridging the representation spaces of speech encoders and LLMs via a small connector module, with a focus on fully…