33 citations · 80 across the 8 of their papers we have counts for
17 papers
Not Quite My Tempo: Voice Activity-aware Speech Synthesis for Lip-Synchronous Dubbing
Alejandro Pérez-González-de-Martos, Florian Lux, Angelina Elizarova +3
Automatic lip-synchronous dubbing requires a speech synthesis model to generate alternating voice and silence patterns in the target language that match the timing of the source cl…
Take the Hint: Improving Arabic Diacritization with Partially-Diacritized Text
Parnia Bahar, Mattia Di Gangi, Nick Rossenbach +1
Automatic Arabic diacritization is useful in many applications, ranging from reading support for language learners to accurate pronunciation predictor for downstream tasks like spe…
On Knowledge Distillation for Direct Speech Translation
Marco Gaido, Mattia A. Di Gangi, Matteo Negri +1
Direct speech translation (ST) has shown to be a complex task requiring knowledge transfer from its sub-tasks: automatic speech recognition (ASR) and machine translation (MT). For…
On Target Segmentation for Direct Speech Translation
Mattia Antonino Di Gangi, Marco Gaido, Matteo Negri +1
Recent studies on direct speech translation show continuous improvements by means of data augmentation techniques and bigger deep learning models. While these methods are helping t…
Contextualized Translation of Automatically Segmented Speech
Marco Gaido, Mattia Antonino Di Gangi, Matteo Negri +2
Direct speech-to-text translation (ST) models are usually trained on corpora segmented at sentence level, but at inference time they are commonly fed with audio split by a voice ac…
Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE Corpus
Luisa Bentivogli, Beatrice Savoldi, Matteo Negri +3
Translating from languages without productive grammatical gender like English into gender-marked languages is a well-known difficulty for machines. This difficulty is also due to t…