activity
20192022
most citedMulti-hop Reading Comprehension across Multiple Documents by Reasoning over Heterogeneous Graphs

28 citations · 39 across the 7 of their papers we have counts for

collaborators

10 papers

cs.CL2022

Unified Speech-Text Pre-training for Speech Translation and Recognition

Yun Tang, Hongyu Gong, Ning Dong +8

We describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method incorporates four sel…

cs.CL2021

FST: the FAIR Speech Translation System for the IWSLT21 Multilingual Shared Task

Yun Tang, Hongyu Gong, Xian Li +4

In this paper, we describe our end-to-end multilingual speech translation system submitted to the IWSLT 2021 evaluation campaign on the Multilingual Speech Translation shared task.…

cs.CL20212 cited

Improving Speech Translation by Understanding and Learning from the Auxiliary Text Translation Task

Yun Tang, Juan Pino, Xian Li +2

Pretraining and multitask learning are widely used to improve the speech to text translation performance. In this study, we are interested in training a speech to text translation…

cs.CL20215 cited

Pay Better Attention to Attention: Head Selection in Multilingual and Multi-Domain Sequence Modeling

Hongyu Gong, Yun Tang, Juan Pino +1

Multi-head attention has each of the attention heads collect salient information from different parts of an input sequence, making it a powerful mechanism for sequence modeling. Mu…

cs.CL2020

A General Multi-Task Learning Framework to Leverage Text Data for Speech to Text Tasks

Yun Tang, Juan Pino, Changhan Wang +2

Attention-based sequence-to-sequence modeling provides a powerful and elegant solution for applications that need to map one sequence to a different sequence. Its success heavily r…

cs.CL2020

Self-Training for End-to-End Speech Translation

Juan Pino, Qiantong Xu, Xutai Ma +2

One of the main challenges for end-to-end speech translation is data scarcity. We leverage pseudo-labels generated from unlabeled audio by a cascade and an end-to-end speech transl…