papers

Publications (111)

cond-mat.stat-mech2025

Pseudo-likelihood produces associative memories able to generalize, even for asymmetric couplings

Francesco D'Amico, Dario Bocchi, Luca Maria Del Bono +2

Energy-based probabilistic models learned by maximizing the likelihood of the data are limited by the intractability of the partition function. A widely used workaround is to maxim…

cs.CL2023

Joint Speech Translation and Named Entity Recognition

Marco Gaido, Sara Papi, Matteo Negri +1

Modern automatic translation systems aim at place the human at the center by providing contextual support and knowledge. In this context, a critical task is enriching the output wi…

cs.CL2021

The Multilingual TEDx Corpus for Speech Recognition and Translation

Elizabeth Salesky, Matthew Wiesner, Jacob Bremerman +5

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a co…

math.AP2025

Energy release and Griffith's criterion for phase-field fracture

Eleonora Maggiorelli, Matteo Negri

Phase field evolutions are obtained by means of time discrete schemes, providing (or selecting) at each time step an equilibrium configuration of the system, which is usually compu…

cs.CL2024

How do Hyenas deal with Human Speech? Speech Recognition and Translation with ConfHyena

Marco Gaido, Sara Papi, Matteo Negri +1

The attention mechanism, a cornerstone of state-of-the-art neural models, faces computational hurdles in processing long sequences due to its quadratic complexity. Consequently, re…

cs.CL2023

Direct Models for Simultaneous Translation and Automatic Subtitling: FBK@IWSLT2023

Sara Papi, Marco Gaido, Matteo Negri

This paper describes the FBK's participation in the Simultaneous Translation and Automatic Subtitling tracks of the IWSLT 2023 Evaluation Campaign. Our submission focused on the us…

cs.CL2020

End-to-End Speech-Translation with Knowledge Distillation: FBK@IWSLT2020

Marco Gaido, Mattia Antonino Di Gangi, Matteo Negri +1

This paper describes FBK's participation in the IWSLT 2020 offline speech translation (ST) task. The task evaluates systems' ability to translate English TED talks audio into Germa…

cs.CL2025

Challenging the Abilities of Large Language Models in Italian: a Community Initiative

Malvina Nissim, Danilo Croce, Viviana Patti +78

The rapid progress of Large Language Models (LLMs) has transformed natural language processing and broadened its impact across research and society. Yet, systematic evaluation of t…

cs.CL2021

How to Split: the Effect of Word Segmentation on Gender Bias in Speech Translation

Marco Gaido, Beatrice Savoldi, Luisa Bentivogli +2

Having recognized gender bias as a major issue affecting current translation technologies, researchers have primarily attempted to mitigate it by working on the data front. However…

cs.LG2020

Wide flat minima and optimal generalization in classifying high-dimensional Gaussian mixtures

Carlo Baldassi, Enrico M. Malatesta, Matteo Negri +1

We analyze the connection between minimizers with good generalizing properties and high local entropy regions of a threshold-linear classifier in Gaussian mixtures with the mean sq…

cs.CL2026

AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech Translation

Sara Papi, Marco Turchi, Matteo Negri

Attention is the core mechanism of today's most used architectures for natural language processing and has been analyzed from many perspectives, including its effectiveness for mac…

cs.LG2026

Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data

Bao Pham, Mohammed J. Zaki, Luca Ambrogioni +2

When do language diffusion models memorize their training data, and how to quantitatively assess their true generative regime? We address these questions by showing that Uniform-ba…

cs.CL2023

No Pitch Left Behind: Addressing Gender Unbalance in Automatic Speech Recognition through Pitch Manipulation

Dennis Fucci, Marco Gaido, Matteo Negri +2

Automatic speech recognition (ASR) systems are known to be sensitive to the sociolinguistic variability of speech data, in which gender plays a crucial role. This can result in dis…

cs.CL2021

CTC-based Compression for Direct Speech Translation

Marco Gaido, Mauro Cettolo, Matteo Negri +1

Previous studies demonstrated that a dynamic phone-informed compression of the input audio is beneficial for speech translation (ST). However, they required a dedicated model for p…

math.AP2019

A quasi-static model for craquelure patterns

Matteo Negri

We consider the quasi-static evolution of a brittle layer on a stiff substrate; adhesion between layers is assumed to be elastic. Employing a phase-field approach we obtain the qua…

q-bio.BM2018

Spontaneous domain formation in disordered copolymers as a mechanism for chromosome structuring

Matteo Negri, Marco Gherardi, Guido Tiana +1

Motivated by the problem of domain formation in chromosomes, we studied a co--polymer model where only a subset of the monomers feel attractive interactions. These monomers are dis…

math.AP2021

Homogenization of Griffith's Criterion for brittle Laminates

Matteo Negri

We consider a periodic, linear elastic laminate with a brittle crack, evolving along a prescribed path according to Griffith's criterion. We study the homogenized limit of this evo…

cs.CL2026

Cross-Attention is Half Explanation in Speech-to-Text Models

Sara Papi, Dennis Fucci, Marco Gaido +2

Cross-attention is a core mechanism in encoder-decoder architectures, widespread in many fields, including speech-to-text (S2T) processing. Its scores have been repurposed for vari…

cs.CL2026

Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems

Marco Gaido, Sara Papi, Mauro Cettolo +2

Streaming Speech-to-Text Translation (StreamST) requires producing translations concurrently with incoming speech under strict latency constraints, demanding models that balance lo…

cs.CL2022

Does Simultaneous Speech Translation need Simultaneous Models?

Sara Papi, Marco Gaido, Matteo Negri +1

In simultaneous speech translation (SimulST), finding the best trade-off between high translation quality and low latency is a challenging task. To meet the latency constraints pos…

cond-mat.dis-nn2023

Storage and Learning phase transitions in the Random-Features Hopfield Model

Matteo Negri, Clarissa Lauditi, Gabriele Perugini +2

The Hopfield model is a paradigmatic model of neural networks that has been analyzed for many decades in the statistical physics, neuroscience, and machine learning communities. In…

cs.CL2020

Contextualized Translation of Automatically Segmented Speech

Marco Gaido, Mattia Antonino Di Gangi, Matteo Negri +2

Direct speech-to-text translation (ST) models are usually trained on corpora segmented at sentence level, but at inference time they are commonly fed with audio split by a voice ac…

cs.CL2025

SPES: Spectrogram Perturbation for Explainable Speech-to-Text Generation

Dennis Fucci, Marco Gaido, Beatrice Savoldi +3

Spurred by the demand for interpretable models, research on eXplainable AI for language technologies has experienced significant growth, with feature attribution methods emerging a…

cs.CL2024

Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?

Marco Gaido, Sara Papi, Matteo Negri +1

The field of natural language processing (NLP) has recently witnessed a transformative shift with the emergence of foundation models, particularly Large Language Models (LLMs) that…

cs.CL2020

Low Resource Neural Machine Translation: A Benchmark for Five African Languages

Surafel M. Lakew, Matteo Negri, Marco Turchi

Recent advents in Neural Machine Translation (NMT) have shown improvements in low-resource language (LRL) translation tasks. In this work, we benchmark NMT between English and five…

cs.CL2026

RedVox: Safety and Fairness Gaps in Speech Models Across Languages

Beatrice Savoldi, Sara Papi, Wafa Aissa +2

Speech-capable models are increasingly deployed in real-world applications across languages. Yet their safety and fairness beyond English settings and under naturalistic conditions…

cs.CL2025

Translation in the Hands of Many:Centering Lay Users in Machine Translation Interactions

Beatrice Savoldi, Alan Ramponi, Matteo Negri +1

Converging societal and technical factors have transformed language technologies into user-facing applications used by the general public across languages. Machine Translation (MT)…

cs.CL2024

SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation

Sara Papi, Marco Gaido, Matteo Negri +1

This paper describes the FBK's participation in the Simultaneous Translation Evaluation Campaign at IWSLT 2024. For this year's submission in the speech-to-text translation (ST) su…

eess.AS2018

Fine-tuning on Clean Data for End-to-End Speech Translation: FBK @ IWSLT 2018

Mattia Antonino Di Gangi, Roberto Dessì, Roldano Cattoni +2

This paper describes FBK's submission to the end-to-end English-German speech translation task at IWSLT 2018. Our system relies on a state-of-the-art model based on LSTMs and CNNs,…

cs.CL2024

What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered Study

Beatrice Savoldi, Sara Papi, Matteo Negri +2

Gender bias in machine translation (MT) is recognized as an issue that can harm people and society. And yet, advancements in the field rarely involve people, the final MT users, or…

cs.CL2024

GFG -- Gender-Fair Generation: A CALAMITA Challenge

Simona Frenda, Andrea Piergentili, Beatrice Savoldi +7

Gender-fair language aims at promoting gender equality by using terms and expressions that include all identities and avoid reinforcing gender stereotypes. Implementing gender-fair…

cs.CL2018

Transfer Learning in Multilingual Neural Machine Translation with Dynamic Vocabulary

Surafel M. Lakew, Aliia Erofeeva, Matteo Negri +2

We propose a method to transfer knowledge across neural machine translation (NMT) models by means of a shared dynamic vocabulary. Our approach allows to extend an initial model for…

cs.CL2024

Findings of the IWSLT 2024 Evaluation Campaign

Ibrahim Said Ahmad, Antonios Anastasopoulos, Ondřej Bojar +42

This paper reports on the shared tasks organized by the 21st IWSLT Conference. The shared tasks address 7 scientific challenges in spoken language translation: simultaneous and off…

cs.CL2020

MuST-Cinema: a Speech-to-Subtitles corpus

Alina Karakanta, Matteo Negri, Marco Turchi

Growing needs in localising audiovisual content in multiple languages through subtitles call for the development of automatic solutions for human subtitling. Neural Machine Transla…

cs.CL2026

How to Evaluate Speech Translation with Source-Aware Neural MT Metrics

Mauro Cettolo, Marco Gaido, Matteo Negri +2

Automatic evaluation of ST systems is typically performed by comparing translation hypotheses with one or more reference translations. While effective to some extent, this approach…

cs.CL2020

Breeding Gender-aware Direct Speech Translation Systems

Marco Gaido, Beatrice Savoldi, Luisa Bentivogli +2

In automatic speech translation (ST), traditional cascade approaches involving separate transcription and translation steps are giving ground to increasingly competitive and more r…

cs.CL2019

Adapting Multilingual Neural Machine Translation to Unseen Languages

Surafel M. Lakew, Alina Karakanta, Marcello Federico +2

Multilingual Neural Machine Translation (MNMT) for low-resource languages (LRL) can be enhanced by the presence of related high-resource languages (HRL), but the relatedness of HRL…

cs.CL2018

Improving Zero-Shot Translation of Low-Resource Languages

Surafel M. Lakew, Quintino F. Lotito, Matteo Negri +2

Recent work on multilingual neural machine translation reported competitive performance with respect to bilingual models and surprisingly good performance even on (zeroshot) transl…

math.NA2025

Gamma convergence for a phase-field cohesive energy

Eleonora Maggiorelli, Matteo Negri, Francesco Vicentini +1

Reproducing the key features of fracture behavior under multiaxial stress states is essential for accurate modeling. Experimental evidence indicates that three intrinsic material p…

cs.CL2025

An LLM-as-a-judge Approach for Scalable Gender-Neutral Translation Evaluation

Andrea Piergentili, Beatrice Savoldi, Matteo Negri +1

Gender-neutral translation (GNT) aims to avoid expressing the gender of human referents when the source text lacks explicit cues about the gender of those referents. Evaluating GNT…

math.AP2020

Existence, energy identity and higher time regularity of solutions to a dynamic visco-elastic cohesive interface model

Matteo Negri, Riccardo Scala

We study the dynamics of visco-elastic materials coupled by a common cohesive interface (or, equivalently, {two single domains separated by} a prescribed cohesive crack) in the ant…

cs.CL2025

Gender-Neutral Rewriting in Italian: Models, Approaches, and Trade-offs

Andrea Piergentili, Beatrice Savoldi, Matteo Negri +1

Gender-neutral rewriting (GNR) aims to reformulate text to eliminate unnecessary gender specifications while preserving meaning, a particularly challenging task in grammatical-gend…

cs.CL2022

Efficient yet Competitive Speech Translation: FBK@IWSLT2022

Marco Gaido, Sara Papi, Dennis Fucci +3

The primary goal of this FBK's systems submission to the IWSLT 2022 offline and simultaneous speech translation tasks is to reduce model training costs without sacrificing translat…

cs.CL2017

Automatic Quality Estimation for ASR System Combination

Shahab Jalalvand, Matteo Negri, Daniele Falavigna +2

Recognizer Output Voting Error Reduction (ROVER) has been widely used for system combination in automatic speech recognition (ASR). In order to select the most appropriate words to…

cs.CL2021

Speechformer: Reducing Information Loss in Direct Speech Translation

Sara Papi, Marco Gaido, Matteo Negri +1

Transformer-based models have gained increasing popularity achieving state-of-the-art performance in many research fields including speech translation. However, Transformer's quadr…

physics.optics2025

Low-power multi-mode fiber projector overcomes shallow neural networks classifiers

Daniele Ancora, Matteo Negri, Antonio Gianfrate +5

In the domain of disordered photonics, the characterization of optically opaque materials for light manipulation and imaging is a primary aim. Among various complex devices, multi-…

cs.CL2021

Between Flexibility and Consistency: Joint Generation of Captions and Subtitles

Alina Karakanta, Marco Gaido, Matteo Negri +1

Speech translation (ST) has lately received growing interest for the generation of subtitles without the need for an intermediate source language transcription and timing (i.e. cap…

cond-mat.dis-nn2025

Statistical mechanics of vector Hopfield network near and above saturation

Flavio Nicoletti, Francesco D'Amico, Matteo Negri

We study analytically and numerically a Hopfield fully-connected network with -dimensional vector spins. These networks are models of associative memory that generalize the stan…

cs.CL2025

The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence

Marco Gaido, Sara Papi, Luisa Bentivogli +6

Training large-scale models presents challenges not only in terms of resource requirements but also in terms of their convergence. For this reason, the learning rate (LR) is often…

cs.CL2026

Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios

Giuseppe Attanasio, Beatrice Savoldi, Daniel Chechelnitsky +4

Speech translation (ST) is increasingly adopted in user applications, yet its evaluation largely focuses on decontextualized testbeds and holistic quality, rather than end users' c…

math.AP2019

Weak solutions for gradient flows under monotonicity constraints

Matteo Negri, Masato Kimura

We consider the gradient flow of a quadratic non-autonomous energy under monotonicity constraint in time and natural regularity assumptions. We provide first a notion of weak solut…

cs.LG2026

Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks

Francesco D'Amico, Dario Bocchi, Matteo Negri

Scaling laws in deep learning -- empirical power-law relationships linking model performance to resource growth -- have emerged as simple yet striking regularities across architect…

cs.CL2025

FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian

Sara Papi, Marco Gaido, Luisa Bentivogli +6

The development of speech foundation models (SFMs) like Whisper and SeamlessM4T has significantly advanced the field of speech processing. However, their closed nature--with inacce…

cs.CL2022

Under the Morphosyntactic Lens: A Multifaceted Evaluation of Gender Bias in Speech Translation

Beatrice Savoldi, Marco Gaido, Luisa Bentivogli +2

Gender bias is largely recognized as a problematic phenomenon affecting language technologies, with recent studies underscoring that it might surface differently across languages.…

math.NA2025

AT1 fourth-order isogeometric phase-field modeling of brittle fracture

Luigi Greco, Eleonora Maggiorelli, Matteo Negri +2

A crucial aspect in phase-field modeling, based on the variational formulation of brittle fracture, is the accurate representation of how the fracture surface energy is dissipated…

cond-mat.dis-nn2026

Benchmarking Graph Neural Networks in Solving Hard Constraint Satisfaction Problems

Geri Skenderi, Lorenzo Buffoni, Francesco D'Amico +6

Graph neural networks (GNNs) are increasingly applied to hard optimization problems, often claiming superiority over classical heuristics. However, such claims risk being unsolid d…

cs.CL2021

Cascade versus Direct Speech Translation: Do the Differences Still Make a Difference?

Luisa Bentivogli, Mauro Cettolo, Marco Gaido +4

Five years after the first published proofs of concept, direct approaches to speech translation (ST) are now competing with traditional cascade solutions. In light of this steady p…

cs.CL2020

Is 42 the Answer to Everything in Subtitling-oriented Speech Translation?

Alina Karakanta, Matteo Negri, Marco Turchi

Subtitling is becoming increasingly important for disseminating information, given the enormous amounts of audiovisual content becoming available daily. Although Neural Machine Tra…

cs.CL2025

The Unheard Alternative: Contrastive Explanations for Speech-to-Text Models

Lina Conti, Dennis Fucci, Marco Gaido +3

Contrastive explanations, which indicate why an AI system produced one output (the target) instead of another (the foil), are widely regarded in explainable AI as more informative…

cs.CL2024

Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection

Beomseok Lee, Marco Gaido, Ioan Calapodescu +2

While crowdsourcing is an established solution for facilitating and scaling the collection of speech data, the involvement of non-experts necessitates protocols to ensure final dat…

cs.SD2021

Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct Speech Translation

Marco Gaido, Matteo Negri, Mauro Cettolo +1

The audio segmentation mismatch between training data and those seen at run-time is a major problem in direct speech translation. Indeed, while systems are usually trained on manua…

cs.CL2019

Machine Translation for Machines: the Sentiment Classification Use Case

Amirhossein Tebbifakhr, Luisa Bentivogli, Matteo Negri +1

We propose a neural machine translation (NMT) approach that, instead of pursuing adequacy and fluency ("human-oriented" quality criteria), aims to generate translations that are be…

cs.CL2023

How To Build Competitive Multi-gender Speech Translation Models For Controlling Speaker Gender Translation

Marco Gaido, Dennis Fucci, Matteo Negri +1

When translating from notional gender languages (e.g., English) into grammatical gender languages (e.g., Italian), the generated translation requires explicit gender assignments fo…

cs.LG2026

Memorization to Generalization: Emergence of Diffusion Models from Associative Memory

Bao Pham, Gabriel Raya, Matteo Negri +3

Dense Associative Memories (DenseAMs) are generalizations of Hopfield networks, which have superior information storage capacity and can store training data points (memories) at lo…

cs.CL2024

When Good and Reproducible Results are a Giant with Feet of Clay: The Importance of Software Quality in NLP

Sara Papi, Marco Gaido, Andrea Pilzer +1

Despite its crucial role in research experiments, code correctness is often presumed only on the basis of the perceived quality of results. This assumption comes with the risk of e…

cs.LG2024

Self-attention as an attractor network: transient memories without backpropagation

Francesco D'Amico, Matteo Negri

Transformers are one of the most successful architectures of modern neural networks. At their core there is the so-called attention mechanism, which recently interested the physics…

math.AP2019

Analysis of staggered evolutions for nonlinear energies in phase field fracture

Stefano Almi, Matteo Negri

We consider a class of separately convex phase field energies employed in fracture mechanics, featuring non-interpenetration and a general softening behavior. We analyze the time-d…

cs.CL2021

Self-Learning for Zero Shot Neural Machine Translation

Surafel M. Lakew, Matteo Negri, Marco Turchi

Neural Machine Translation (NMT) approaches employing monolingual data are showing steady improvements in resource rich conditions. However, evaluations using real-world low-resour…

cs.SD2024

StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection

Sara Papi, Marco Gaido, Matteo Negri +1

Streaming speech-to-text translation (StreamST) is the task of automatically translating speech while incrementally receiving an audio stream. Unlike simultaneous ST (SimulST), whi…

cs.CL2023

Gender Neutralization for an Inclusive Machine Translation: from Theoretical Foundations to Open Challenges

Andrea Piergentili, Dennis Fucci, Beatrice Savoldi +2

Gender inclusivity in language technologies has become a prominent research topic. In this study, we explore gender-neutral translation (GNT) as a form of gender inclusivity and a…

cs.CL2023

Hi Guys or Hi Folks? Benchmarking Gender-Neutral Machine Translation with the GeNTE Corpus

Andrea Piergentili, Beatrice Savoldi, Dennis Fucci +2

Gender inequality is embedded in our communication practices and perpetuated in translation technologies. This becomes particularly apparent when translating into grammatical gende…

eess.AS2026

SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation

Amirbek Djanibekov, Luisa Bentivogli, Matteo Negri +1

Simultaneous speech-to-speech translation (SimulS2S) is essential for real-time multilingual communication, with increasing integration into meeting and streaming platforms. Despit…

cs.CL2023

Direct Speech Translation for Automatic Subtitling

Sara Papi, Marco Gaido, Alina Karakanta +3

Automatic subtitling is the task of automatically translating the speech of audiovisual content into short pieces of timed text, i.e. subtitles and their corresponding timestamps.…

q-bio.BM2021

Native state of natural proteins optimises local entropy

Matteo Negri, Guido Tiana, Riccardo Zecchina

The differing ability of polypeptide conformations to act as the native state of proteins has long been rationalized in terms of differing kinetic accessibility or thermodynamic st…

cs.CL2020

On Target Segmentation for Direct Speech Translation

Mattia Antonino Di Gangi, Marco Gaido, Matteo Negri +1

Recent studies on direct speech translation show continuous improvements by means of data augmentation techniques and bigger deep learning models. While these methods are helping t…

cs.CL2023

Test Suites Task: Evaluation of Gender Fairness in MT with MuST-SHE and INES

Beatrice Savoldi, Marco Gaido, Matteo Negri +1

As part of the WMT-2023 "Test suites" shared task, in this paper we summarize the results of two test suites evaluations: MuST-SHE-WMT23 and INES. By focusing on the en-de and de-e…

cs.CL2024

Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond

Beomseok Lee, Ioan Calapodescu, Marco Gaido +2

We present Speech-MASSIVE, a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE…

cond-mat.dis-nn2024

Daydreaming Hopfield Networks and their surprising effectiveness on correlated data

Ludovica Serricchio, Dario Bocchi, Claudio Chilin +4

To improve the storage capacity of the Hopfield model, we develop a version of the dreaming algorithm that perpetually reinforces the patterns to be stored (as in the Hebb rule), a…

cs.CL2024

MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages

Marco Gaido, Sara Papi, Luisa Bentivogli +6

The rise of foundation models (FMs), coupled with regulatory efforts addressing their risks and impacts, has sparked significant interest in open-source models. However, existing s…

cs.CL2023

Integrating Language Models into Direct Speech Translation: An Inference-Time Solution to Control Gender Inflection

Dennis Fucci, Marco Gaido, Sara Papi +3

When translating words referring to the speaker, speech translation (ST) systems should not resort to default masculine generics nor rely on potentially misleading vocal traits. Ra…

cs.CL2026

Generative AI Practices, Literacy, and Divides: An Empirical Analysis in the Italian Context

Beatrice Savoldi, Giuseppe Attanasio, Olga Gorodetskaya +10

The rise of generative AI (GenAI) chatbots accessible via conversational interfaces is transforming digital interactions and holds economic promise. However, these tools might deep…

cs.CL2024

SBAAM! Eliminating Transcript Dependency in Automatic Subtitling

Marco Gaido, Sara Papi, Matteo Negri +2

Subtitling plays a crucial role in enhancing the accessibility of audiovisual content and encompasses three primary subtasks: translating spoken dialogue, segmenting translations i…

cs.CL2017

Linguistically Motivated Vocabulary Reduction for Neural Machine Translation from Turkish to English

Duygu Ataman, Matteo Negri, Marco Turchi +1

The necessity of using a fixed-size word vocabulary in order to control the model complexity in state-of-the-art neural machine translation (NMT) systems is an important bottleneck…

cs.CL2025

Different Speech Translation Models Encode and Translate Speaker Gender Differently

Dennis Fucci, Marco Gaido, Matteo Negri +3

Recent studies on interpreting the hidden states of speech models have shown their ability to capture speaker-specific features, including gender. Does this finding also hold for s…

cs.CL2021

Dealing with training and test segmentation mismatch: FBK@IWSLT2021

Sara Papi, Marco Gaido, Matteo Negri +1

This paper describes FBK's system submission to the IWSLT 2021 Offline Speech Translation task. We participated with a direct model, which is a Transformer-based architecture train…

cs.CL2022

Dodging the Data Bottleneck: Automatic Subtitling with Automatically Segmented ST Corpora

Sara Papi, Alina Karakanta, Matteo Negri +1

Speech translation for subtitling (SubST) is the task of automatically translating speech data into well-formed subtitles by inserting subtitle breaks compliant to specific display…

cs.CL2020

On Knowledge Distillation for Direct Speech Translation

Marco Gaido, Mattia A. Di Gangi, Matteo Negri +1

Direct speech translation (ST) has shown to be a complex task requiring knowledge transfer from its sub-tasks: automatic speech recognition (ASR) and machine translation (MT). For…

cs.CL2025

Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTE

Beatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin +7

Avoiding the propagation of undue (binary) gender inferences and default masculine language remains a key challenge towards inclusive multilingual technologies, particularly when t…

cs.CL2021

Simultaneous Speech Translation for Live Subtitling: from Delay to Display

Alina Karakanta, Sara Papi, Matteo Negri +1

With the increased audiovisualisation of communication, the need for live subtitles in multilingual events is more relevant than ever. In an attempt to automatise the process, we a…

cs.CL2020

Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE Corpus

Luisa Bentivogli, Beatrice Savoldi, Matteo Negri +3

Translating from languages without productive grammatical gender like English into gender-marked languages is a well-known difficulty for machines. This difficulty is also due to t…

cs.CL2026

FBK's Long-form SpeechLLMs for IWSLT 2026 Instruction Following

Zhihang Xie, Marco Gaido, Sara Papi +2

This paper describes our submission to the IWSLT 2026 Instruction Following shared task. SpeechLLMs are developed for both short-form and long-form speech instruction following und…

cs.CL2024

Enhancing Gender-Inclusive Machine Translation with Neomorphemes and Large Language Models

Andrea Piergentili, Beatrice Savoldi, Matteo Negri +1

Machine translation (MT) models are known to suffer from gender bias, especially when translating into languages with extensive gendered morphology. Accordingly, they still fall sh…

cs.CL2022

Over-Generation Cannot Be Rewarded: Length-Adaptive Average Lagging for Simultaneous Speech Translation

Sara Papi, Marco Gaido, Matteo Negri +1

Simultaneous speech translation (SimulST) systems aim at generating their output with the lowest possible latency, which is normally computed in terms of Average Lagging (AL). In t…

q-bio.QM2019

Natural representation of composite data with replicated autoencoders

Matteo Negri, Davide Bergamini, Carlo Baldassi +2

Generative processes in biology and other fields often produce data that can be regarded as resulting from a composition of basic features. Here we present an unsupervised method b…

cond-mat.dis-nn2024

Random Features Hopfield Networks generalize retrieval to previously unseen examples

Silvio Kalaj, Clarissa Lauditi, Gabriele Perugini +3

It has been recently shown that a learning transition happens when a Hopfield Network stores examples generated as superpositions of random features, where new attractors correspon…

cs.CL2021

Is "moby dick" a Whale or a Bird? Named Entities and Terminology in Speech Translation

Marco Gaido, Susana Rodríguez, Matteo Negri +2

Automatic translation systems are known to struggle with rare words. Among these, named entities (NEs) and domain-specific terms are crucial, since errors in their translation can…

cs.CL2021

Gender Bias in Machine Translation

Beatrice Savoldi, Marco Gaido, Luisa Bentivogli +2

Machine translation (MT) technology has facilitated our daily tasks by providing accessible shortcuts for gathering, elaborating and communicating information. However, it can suff…

cs.CL2017

DNN adaptation by automatic quality estimation of ASR hypotheses

Daniele Falavigna, Marco Matassoni, Shahab Jalalvand +2

In this paper we propose to exploit the automatic Quality Estimation (QE) of ASR hypotheses to perform the unsupervised adaptation of a deep neural network modeling acoustic probab…

cs.CL2024

Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation

Matthias Sperber, Ondřej Bojar, Barry Haddow +8

Human evaluation is a critical component in machine translation system development and has received much attention in text translation research. However, little prior work exists o…

cs.CL2026

Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation

Lina Conti, Dennis Fucci, Marco Gaido +3

Unlike text, speech conveys information about the speaker, such as gender, through acoustic cues like pitch. This gives rise to modality-specific bias concerns. For example, in spe…