2 citations · 3 across the 4 of their papers we have counts for
4 papers
End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation
Juan Zuluaga-Gomez, Zhaocheng Huang, Xing Niu +5
Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversat…
Speaker Diarization of Scripted Audiovisual Content
Yogesh Virkar, Brian Thompson, Rohit Paturi +2
The media localization industry usually requires a verbatim script of the final film or TV production in order to create subtitles or dubbing scripts in a foreign language. In part…
Improving Isochronous Machine Translation with Target Factors and Auxiliary Counters
Proyag Pal, Brian Thompson, Yogesh Virkar +3
To translate speech for automatic dubbing, machine translation needs to be isochronous, i.e. translated speech needs to be aligned with the source in terms of speech durations. We…
Jointly Optimizing Translations and Speech Timing to Improve Isochrony in Automatic Dubbing
Alexandra Chronopoulou, Brian Thompson, Prashant Mathur +3
Automatic dubbing (AD) is the task of translating the original speech in a video into target language speech. The new target language speech should satisfy isochrony; that is, the…