activity
20172022
most citedWikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia

67 citations · 102 across the 15 of their papers we have counts for

collaborators

18 papers

cs.CL20224 cited

SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations

Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong +7

We present SpeechMatrix, a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. It contains speech alignments…

cs.CL2022

Unified Speech-Text Pre-training for Speech Translation and Recognition

Yun Tang, Hongyu Gong, Ning Dong +8

We describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method incorporates four sel…

cs.CL2021

FST: the FAIR Speech Translation System for the IWSLT21 Multilingual Shared Task

Yun Tang, Hongyu Gong, Xian Li +4

In this paper, we describe our end-to-end multilingual speech translation system submitted to the IWSLT 2021 evaluation campaign on the Multilingual Speech Translation shared task.…

cs.CL20215 cited

Pay Better Attention to Attention: Head Selection in Multilingual and Multi-Domain Sequence Modeling

Hongyu Gong, Yun Tang, Juan Pino +1

Multi-head attention has each of the attention heads collect salient information from different parts of an input sequence, making it a powerful mechanism for sequence modeling. Mu…

cs.CL20213 cited

LAWDR: Language-Agnostic Weighted Document Representations from Pre-trained Models

Hongyu Gong, Vishrav Chaudhary, Yuqing Tang +1

Cross-lingual document representations enable language understanding in multilingual contexts and allow transfer learning from high-resource to low-resource languages at the docume…

cs.CL20212 cited

Abusive Language Detection in Heterogeneous Contexts: Dataset Collection and the Role of Supervised Attention

Hongyu Gong, Alberto Valido, Katherine M. Ingram +3

Abusive language is a massive problem in online social platforms. Existing abusive language detection techniques are particularly ill-suited to comments containing heterogeneous ab…