6 papers
grapheme-kit: Grapheme-Level Metrics and Text Processing for Multilingual NLP
Izzath Nisfer, Ashini Kavindya, Ovindu Atukorala +2
Existing lexical distance, similarity, and evaluation metrics operate on Unicode code points, which can misrepresent errors in writing systems where a single grapheme is represente…
emb-diversity: A Tool for Embedding-Based Measurement of Data Diversity
Cantao Su, Menan Velayuthan, Esther Ploeger +2
There is growing evidence that data diversity is crucial for developing fair and robust NLP models. However, current approaches to measure diversity remain inconsistent and fragmen…
Confusion-Aware Transfer Teacher Curriculum Learning Framework: Disentangling Scoring and Pacing Effects
Savini Kommalage, Sanka Mohottala, Asiri Gawesha +5
Curriculum learning couples two design choices, how samples are scored by difficulty and how harder samples are paced into training, making it difficult to attribute observed gains…
From Phonemes to Meaning: Evaluating Large Language Models on Tamil
Jeyarajalingam Varsha, Menan Velayuthan, Sumirtha Karunakaran +2
Large Language Models (LLMs) have shown strong generalization across tasks in high-resource languages; however, their linguistic competence in low-resource and morphologically rich…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
Egalitarian Language Representation in Language Models: It All Begins with Tokenizers
Menan Velayuthan, Kengatharaiyer Sarveswaran
Tokenizers act as a bridge between human language and the latent space of language models, influencing how language is represented in these models. Due to the immense popularity of…