21 papers
Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems
Marco Gaido, Sara Papi, Mauro Cettolo +2
Streaming Speech-to-Text Translation (StreamST) requires producing translations concurrently with incoming speech under strict latency constraints, demanding models that balance lo…
RedVox: Safety and Fairness Gaps in Speech Models Across Languages
Beatrice Savoldi, Sara Papi, Wafa Aissa +2
Speech-capable models are increasingly deployed in real-world applications across languages. Yet their safety and fairness beyond English settings and under naturalistic conditions…
FBK's Long-form SpeechLLMs for IWSLT 2026 Instruction Following
Zhihang Xie, Marco Gaido, Sara Papi +2
This paper describes our submission to the IWSLT 2026 Instruction Following shared task. SpeechLLMs are developed for both short-form and long-form speech instruction following und…
Cross-Attention is Half Explanation in Speech-to-Text Models
Sara Papi, Dennis Fucci, Marco Gaido +2
Cross-attention is a core mechanism in encoder-decoder architectures, widespread in many fields, including speech-to-text (S2T) processing. Its scores have been repurposed for vari…
Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios
Giuseppe Attanasio, Beatrice Savoldi, Daniel Chechelnitsky +4
Speech translation (ST) is increasingly adopted in user applications, yet its evaluation largely focuses on decontextualized testbeds and holistic quality, rather than end users' c…
Generative AI Practices, Literacy, and Divides: An Empirical Analysis in the Italian Context
Beatrice Savoldi, Giuseppe Attanasio, Olga Gorodetskaya +10
The rise of generative AI (GenAI) chatbots accessible via conversational interfaces is transforming digital interactions and holds economic promise. However, these tools might deep…