60 citations · 150 across the 8 of their papers we have counts for
11 papers · 1 filter
ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese
Hai-Chung Nguyen-Phung, Ngoc C. Lê, Van-Chien Nguyen +2
After two years of appearance, COVID-19 has negatively affected people and normal life around the world. As in May 2022, there are more than 522 million cases and six million death…
Speech-to-Speech Translation For A Real-world Unwritten Language
Peng-Jen Chen, Kevin Tran, Yilin Yang +13
We study speech-to-speech translation (S2ST) that translates speech from one language into another language and focuses on building systems to support languages without standard te…
SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations
Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong +7
We present SpeechMatrix, a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. It contains speech alignments…
Larger-Scale Transformers for Multilingual Masked Language Modeling
Naman Goyal, Jingfei Du, Myle Ott +2
Recent work has demonstrated the effectiveness of cross-lingual language model pretraining for cross-lingual understanding. In this study, we present the results of two larger mult…
Supervised Contrastive Learning for Pre-trained Language Model Fine-tuning
Beliz Gunel, Jingfei Du, Alexis Conneau +1
State-of-the-art natural language understanding classification models follow two-stages: pre-training a large language model on an auxiliary task, and then fine-tuning the model on…
Self-training Improves Pre-training for Natural Language Understanding
Jingfei Du, Edouard Grave, Beliz Gunel +5
Unsupervised pre-training has led to much recent progress in natural language understanding. In this paper, we study self-training as another way to leverage unlabeled data through…