4 papers
Rethinking Cross-lingual Gaps from a Statistical Viewpoint
Vihari Piratla, Purvam Jain, Darshan Singh +3
Any piece of knowledge is usually expressed in one or a handful of natural languages on the web or in any large corpus. Large Language Models (LLMs) act as a bridge by acquiring kn…
VAANI: Capturing the language landscape for an inclusive digital India
Sujith Pulikodan, Abhayjeet Singh, Agneedh Basu +17
Voice based technologies have the potential to bridge digital accessibility gaps; however, existing datasets fail to capture the linguistic and regional diversity of Indic language…
SMOL: Professionally translated parallel data for 115 under-represented languages
Isaac Caswell, Elizabeth Nielsen, Jiaming Luo +23
We open-source SMOL (Set of Maximal Overall Leverage), a suite of training data to unlock machine translation for low-resource languages. SMOL has been translated into 124 (and gro…
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
Harman Singh, Nitish Gupta, Shikhar Bharadwaj +2
As large language models (LLMs) see increasing adoption across the globe, it is imperative for LLMs to be representative of the linguistic diversity of the world. India is a lingui…