5 papers
FLiP: Towards understanding and interpreting multimodal multilingual sentence embeddings
Santosh Kesiraju, Bolaji Yusuf, Å imon SedláÄek +2
This paper presents factorized linear projection (FLiP) models for understanding pretrained sentence embedding spaces. We train FLiP models to recover the lexical content from mult…
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
Sonal Kumar, Å imon SedláÄek, Vaibhavi Lokegaonkar +31
Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio unde…
Factors affecting the in-context learning abilities of LLMs for dialogue state tracking
Pradyoth Hegde, Santosh Kesiraju, Jan Å vec +5
This study explores the application of in-context learning (ICL) to the dialogue state tracking (DST) problem and investigates the factors that influence its effectiveness. We use…
Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
Å imon SedláÄek, Bolaji Yusuf, Ján Å vec +4
In this work, we approach spoken Dialogue State Tracking (DST) by bridging the representation spaces of speech encoders and LLMs via a small connector module, with a focus on fully…
Aligning Pre-trained Models for Spoken Language Translation
Å imon SedláÄek, Santosh Kesiraju, Alexander Polok +1
This paper investigates a novel approach to end-to-end speech translation (ST) based on aligning frozen pre-trained automatic speech recognition (ASR) and machine translation (MT)…