activity
20242026
collaborators

5 papers

cs.CL2026

FLiP: Towards understanding and interpreting multimodal multilingual sentence embeddings

Santosh Kesiraju, Bolaji Yusuf, Šimon Sedláček +2

This paper presents factorized linear projection (FLiP) models for understanding pretrained sentence embedding spaces. We train FLiP models to recover the lexical content from mult…

eess.AS2025

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Sonal Kumar, Šimon Sedláček, Vaibhavi Lokegaonkar +31

Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio unde…

cs.CL2025

Factors affecting the in-context learning abilities of LLMs for dialogue state tracking

Pradyoth Hegde, Santosh Kesiraju, Jan Å vec +5

This study explores the application of in-context learning (ICL) to the dialogue state tracking (DST) problem and investigates the factors that influence its effectiveness. We use…

eess.AS2025

Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs

Šimon Sedláček, Bolaji Yusuf, Ján Švec +4

In this work, we approach spoken Dialogue State Tracking (DST) by bridging the representation spaces of speech encoders and LLMs via a small connector module, with a focus on fully…

cs.CL2024

Aligning Pre-trained Models for Spoken Language Translation

Šimon Sedláček, Santosh Kesiraju, Alexander Polok +1

This paper investigates a novel approach to end-to-end speech translation (ST) based on aligning frozen pre-trained automatic speech recognition (ASR) and machine translation (MT)…