activity
20242026
collaborators

10 papers

cs.SD2026

ORCA: Open-ended Response Correctness Assessment for Audio Question Answering

Šimon Sedláček, Sara Barahona, Bolaji Yusuf +9

Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art. As benchmarks rapidly evolve to incorporate complex reas…

cs.CL2026

Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking

Katia Vendrame, Bolaji Yusuf, Santosh Kesiraju +3

End-to-end spoken dialogue state tracking (DST) is made difficult by the tandem of having to handle speech input and data scarcity. Combining speech foundation encoders and large l…

cs.CL2026

FLiP: Towards understanding and interpreting multimodal multilingual sentence embeddings

Santosh Kesiraju, Bolaji Yusuf, Šimon Sedláček +2

This paper presents factorized linear projection (FLiP) models for understanding pretrained sentence embedding spaces. We train FLiP models to recover the lexical content from mult…

eess.AS2025

DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition

Alexander Polok, Santosh Kesiraju, Karel Beneš +3

This paper presents a simple yet effective regularization for the internal language model induced by the decoder in encoder-decoder ASR models, thereby improving robustness and gen…

cs.CL2025

Factors affecting the in-context learning abilities of LLMs for dialogue state tracking

Pradyoth Hegde, Santosh Kesiraju, Jan Å vec +5

This study explores the application of in-context learning (ICL) to the dialogue state tracking (DST) problem and investigates the factors that influence its effectiveness. We use…

eess.AS2025

Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs

Šimon Sedláček, Bolaji Yusuf, Ján Švec +4

In this work, we approach spoken Dialogue State Tracking (DST) by bridging the representation spaces of speech encoders and LLMs via a small connector module, with a focus on fully…