activity
20192022
most citedMuRIL: Multilingual Representations for Indian Languages

160 citations · 276 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL202215 cited

FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

Alexis Conneau, Min Ma, Simran Khanuja +6

We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark. FLEURS is an n-way parallel speech dataset in 102 languages built on top of…

cs.CL20221 cited

XTREME-S: Evaluating Cross-lingual Speech Representations

Alexis Conneau, Ankur Bapna, Yu Zhang +16

We introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages. XTREME-S covers four task families: speech recognition, classif…

cs.CL202259 cited

mSLAM: Massively multilingual joint pre-training for speech and text

Ankur Bapna, Colin Cherry, Yu Zhang +6

We present mSLAM, a multilingual Speech and LAnguage Model that learns cross-lingual cross-modal representations of speech and text by pre-training jointly on large amounts of unla…

cs.CL20211 cited

MergeDistill: Merging Pre-trained Language Models using Distillation

Simran Khanuja, Melvin Johnson, Partha Talukdar

Pre-trained multilingual language models (LMs) have achieved state-of-the-art results in cross-lingual transfer, but they often lead to an inequitable representation of languages d…

cs.CL2021160 cited

MuRIL: Multilingual Representations for Indian Languages

Simran Khanuja, Diksha Bansal, Sarvesh Mehtani +11

India is a multilingual society with 1369 rationalized languages and dialects being spoken across the country (INDIA, 2011). Of these, the 22 scheduled languages have a staggering…

cs.CL20204 cited

Cross-lingual and Multilingual Spoken Term Detection for Low-Resource Indian Languages

Sanket Shah, Satarupa Guha, Simran Khanuja +1

Spoken Term Detection (STD) is the task of searching for words or phrases within audio, given either text or spoken input as a query. In this work, we use state-of-the-art Hindi, T…