4 papers
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
André G. Viveiros, Patrick Fernandes, Saul Santos +7
Despite significant advances in vision-language models (VLMs), most existing work follows an English-centric design process, limiting their effectiveness in multilingual settings.…
From TOWER to SPIRE: Adding the Speech Modality to a Translation-Specialist LLM
Kshitij Ambilduke, Ben Peters, Sonal Sannigrahi +5
We introduce Spire, a speech-augmented language model (LM) capable of both translating and transcribing speech input from English into 10 other languages as well as translating tex…
Instituto de Telecomunicações at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning
Giuseppe Attanasio, Sonal Sannigrahi, Ben Peters +1
This paper presents the IT-IST submission to the IWSLT 2025 Shared Task on Instruction Following Speech Processing. We submit results for the Short Track, i.e., speech recognition,…
Movie Facts and Fibs (MF): A Benchmark for Long Movie Understanding
Emmanouil Zaranis, António Farinhas, Saul Santos +28
Despite recent progress in vision-language models (VLMs), holistic understanding of long-form video content remains a significant challenge, partly due to limitations in current be…