28 papers
Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems
Marco Gaido, Sara Papi, Mauro Cettolo +2
Streaming Speech-to-Text Translation (StreamST) requires producing translations concurrently with incoming speech under strict latency constraints, demanding models that balance lo…
RedVox: Safety and Fairness Gaps in Speech Models Across Languages
Beatrice Savoldi, Sara Papi, Wafa Aissa +2
Speech-capable models are increasingly deployed in real-world applications across languages. Yet their safety and fairness beyond English settings and under naturalistic conditions…
FBK's Long-form SpeechLLMs for IWSLT 2026 Instruction Following
Zhihang Xie, Marco Gaido, Sara Papi +2
This paper describes our submission to the IWSLT 2026 Instruction Following shared task. SpeechLLMs are developed for both short-form and long-form speech instruction following und…
Cross-Attention is Half Explanation in Speech-to-Text Models
Sara Papi, Dennis Fucci, Marco Gaido +2
Cross-attention is a core mechanism in encoder-decoder architectures, widespread in many fields, including speech-to-text (S2T) processing. Its scores have been repurposed for vari…
DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
Sara Papi, Luisa Bentivogli
Simultaneous speech-to-text translation (SimulST) generates translations while speech is still unfolding, requiring a streaming policy that decides when to read and when to write.…
Generative AI Practices, Literacy, and Divides: An Empirical Analysis in the Italian Context
Beatrice Savoldi, Giuseppe Attanasio, Olga Gorodetskaya +10
The rise of generative AI (GenAI) chatbots accessible via conversational interfaces is transforming digital interactions and holds economic promise. However, these tools might deep…