activity
20192025
most citedSupervised Contrastive Learning for Pre-trained Language Model Fine-tuning

60 citations · 150 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

11 papers · 1 filter

cs.CL2025

ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese

Hai-Chung Nguyen-Phung, Ngoc C. Lê, Van-Chien Nguyen +2

After two years of appearance, COVID-19 has negatively affected people and normal life around the world. As in May 2022, there are more than 522 million cases and six million death…

cs.CL20225 cited

Speech-to-Speech Translation For A Real-world Unwritten Language

Peng-Jen Chen, Kevin Tran, Yilin Yang +13

We study speech-to-speech translation (S2ST) that translates speech from one language into another language and focuses on building systems to support languages without standard te…

cs.CL20224 cited

SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations

Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong +7

We present SpeechMatrix, a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. It contains speech alignments…

cs.CL2021

Larger-Scale Transformers for Multilingual Masked Language Modeling

Naman Goyal, Jingfei Du, Myle Ott +2

Recent work has demonstrated the effectiveness of cross-lingual language model pretraining for cross-lingual understanding. In this study, we present the results of two larger mult…

cs.CL202060 cited

Supervised Contrastive Learning for Pre-trained Language Model Fine-tuning

Beliz Gunel, Jingfei Du, Alexis Conneau +1

State-of-the-art natural language understanding classification models follow two-stages: pre-training a large language model on an auxiliary task, and then fine-tuning the model on…

cs.CL202046 cited

Self-training Improves Pre-training for Natural Language Understanding

Jingfei Du, Edouard Grave, Beliz Gunel +5

Unsupervised pre-training has led to much recent progress in natural language understanding. In this paper, we study self-training as another way to leverage unlabeled data through…