activity
20182026
most citedMonotonic Multihead Attention

68 citations · 234 across the 34 of their papers we have counts for

collaborators
Showing 2021 · cs.CLShow all

11 papers · 2 filters

cs.CL2021

XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Arun Babu, Changhan Wang, Andros Tjandra +10

This paper presents XLS-R, a large-scale model for cross-lingual speech representation learning based on wav2vec 2.0. We train models with up to 2B parameters on nearly half a mill…

cs.CL2021★ 1 cited

Textless Speech-to-Speech Translation on Real Data

Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne +8

We present a textless speech-to-speech translation (S2ST) system that can translate speech from one language into another language and can be built without the need of any text dat…

cs.CL2021★ 1 cited

Direct Simultaneous Speech-to-Speech Translation with Variational Monotonic Multihead Attention

Xutai Ma, Hongyu Gong, Danni Liu +6

We present a direct simultaneous speech-to-speech translation (Simul-S2ST) model, Furthermore, the generation of translation is independent from intermediate text representations.…

cs.CL2021

From Start to Finish: Latency Reduction Strategies for Incremental Speech Synthesis in Simultaneous Speech-to-Speech Translation

Danni Liu, Changhan Wang, Hongyu Gong +3

Speech-to-speech translation (S2ST) converts input speech to speech in another language. A challenge of delivering S2ST in real time is the accumulated delay between the translatio…

cs.CL2021

FST: the FAIR Speech Translation System for the IWSLT21 Multilingual Shared Task

Yun Tang, Hongyu Gong, Xian Li +4

In this paper, we describe our end-to-end multilingual speech translation system submitted to the IWSLT 2021 evaluation campaign on the Multilingual Speech Translation shared task.…

cs.CL2021★ 2 cited

Improving Speech Translation by Understanding and Learning from the Auxiliary Text Translation Task

Yun Tang, Juan Pino, Xian Li +2

Pretraining and multitask learning are widely used to improve the speech to text translation performance. In this study, we are interested in training a speech to text translation…