activity
20222025
collaborators

8 papers

eess.AS2025

DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition

Alexander Polok, Santosh Kesiraju, Karel Beneš +3

This paper presents a simple yet effective regularization for the internal language model induced by the decoder in encoder-decoder ASR models, thereby improving robustness and gen…

eess.AS2025

Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing

Junyi Peng, Lin Zhang, Jiangyu Han +5

Although large-scale self-supervised learning (SSL) models like WavLM have achieved state-of-the-art performance in speech processing, their significant size impedes deployment on…

eess.AS2025

BUT System for the MLC-SLM Challenge

Alexander Polok, Jiangyu Han, Dominik Klement +3

We present a two-speaker automatic speech recognition (ASR) system that combines DiCoW -- a diarization-conditioned variant of Whisper -- with DiariZen, a diarization pipeline buil…

cs.CL2025

Factors affecting the in-context learning abilities of LLMs for dialogue state tracking

Pradyoth Hegde, Santosh Kesiraju, Jan Švec +5

This study explores the application of in-context learning (ICL) to the dialogue state tracking (DST) problem and investigates the factors that influence its effectiveness. We use…

eess.AS2025

Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs

Šimon Sedláček, Bolaji Yusuf, Ján Švec +4

In this work, we approach spoken Dialogue State Tracking (DST) by bridging the representation spaces of speech encoders and LLMs via a small connector module, with a focus on fully…

cs.CL2024

Aligning Pre-trained Models for Spoken Language Translation

Šimon Sedláček, Santosh Kesiraju, Alexander Polok +1

This paper investigates a novel approach to end-to-end speech translation (ST) based on aligning frozen pre-trained automatic speech recognition (ASR) and machine translation (MT)…