7 papers
Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems
Marco Gaido, Sara Papi, Mauro Cettolo +2
Streaming Speech-to-Text Translation (StreamST) requires producing translations concurrently with incoming speech under strict latency constraints, demanding models that balance lo…
How to Evaluate Speech Translation with Source-Aware Neural MT Metrics
Mauro Cettolo, Marco Gaido, Matteo Negri +2
Automatic evaluation of ST systems is typically performed by comparing translation hypotheses with one or more reference translations. While effective to some extent, this approach…
Challenging the Abilities of Large Language Models in Italian: a Community Initiative
Malvina Nissim, Danilo Croce, Viviana Patti +78
The rapid progress of Large Language Models (LLMs) has transformed natural language processing and broadened its impact across research and society. Yet, systematic evaluation of t…
Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution
Dennis Fucci, Marco Gaido, Matteo Negri +2
Despite significant advances in ASR, the specific acoustic cues models rely on remain unclear. Prior studies have examined such cues on a limited set of phonemes and outdated model…
FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
Sara Papi, Marco Gaido, Luisa Bentivogli +6
The development of speech foundation models (SFMs) like Whisper and SeamlessM4T has significantly advanced the field of speech processing. However, their closed nature--with inacce…
The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence
Marco Gaido, Sara Papi, Luisa Bentivogli +6
Training large-scale models presents challenges not only in terms of resource requirements but also in terms of their convergence. For this reason, the learning rate (LR) is often…