4 papers
Multilingual Long-Form Speech Instruction Following: KIT's Submission to IWSLT 2026
Enes Yavuz Ugan, Maike Züfle, Yuka Ko +5
With the advent of Large Language Models, single-task and token-based multi-task models have evolved into instruction-based systems that infer task and target language implicitly f…
MUSCAT: MUltilingual, SCientific ConversATion Benchmark
Supriti Sinhamahapatra, Thai-Binh Nguyen, YiÄit OÄuz +3
The goal of multilingual speech technology is to facilitate seamless communication between individuals speaking different languages, creating the experience as though everyone were…
Do Slides Help? Multi-modal Context for Automatic Transcription of Conference Talks
Supriti Sinhamahapatra, Jan Niehues
State-of-the-art (SOTA) Automatic Speech Recognition (ASR) systems primarily rely on acoustic information while disregarding additional multi-modal context. However, visual informa…
How do Multimodal Foundation Models Encode Text and Speech? An Analysis of Cross-Lingual and Cross-Modal Representations
Hyunji Lee, Danni Liu, Supriti Sinhamahapatra +1
Multimodal foundation models aim to create a unified representation space that abstracts away from surface features like language syntax or modality differences. To investigate thi…