activity
20172024
most citedCross-lingual Language Model Pretraining

1.6k citations · 2.4k across the 13 of their papers we have counts for

collaborators
Showing cs.CLShow all

21 papers · 1 filter

cs.CL20241 cited

Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation

Kartik Kartik, Sanjana Soni, Anoop Kunchukuttan +2

The widespread online communication in a modern multilingual world has provided opportunities to blend more than one language (aka code-mixed language) in a single utterance. This…

cs.CL2024

DANSK and DaCy 2.6.0: Domain Generalization of Danish Named Entity Recognition

Kenneth Enevoldsen, Emil Trenckner Jessen, Rebekah Baglini

Named entity recognition is one of the cornerstones of Danish NLP, essential for language technology applications within both industry and research. However, Danish NER is inhibite…

cs.CL20232 cited

Toward Joint Language Modeling for Speech Units and Text

Ju-Chieh Chou, Chung-Ming Chien, Wei-Ning Hsu +5

Speech and text are two major forms of human language. The research community has been focusing on mapping speech to text or vice versa for many years. However, in the field of lan…

cs.CL202215 cited

FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

Alexis Conneau, Min Ma, Simran Khanuja +6

We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark. FLEURS is an n-way parallel speech dataset in 102 languages built on top of…

cs.CL20221 cited

XTREME-S: Evaluating Cross-lingual Speech Representations

Alexis Conneau, Ankur Bapna, Yu Zhang +16

We introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages. XTREME-S covers four task families: speech recognition, classif…

cs.CL202259 cited

mSLAM: Massively multilingual joint pre-training for speech and text

Ankur Bapna, Colin Cherry, Yu Zhang +6

We present mSLAM, a multilingual Speech and LAnguage Model that learns cross-lingual cross-modal representations of speech and text by pre-training jointly on large amounts of unla…