activity
20242026
collaborators

6 papers

cs.CL2026

grapheme-kit: Grapheme-Level Metrics and Text Processing for Multilingual NLP

Izzath Nisfer, Ashini Kavindya, Ovindu Atukorala +2

Existing lexical distance, similarity, and evaluation metrics operate on Unicode code points, which can misrepresent errors in writing systems where a single grapheme is represente…

cs.CL2026

emb-diversity: A Tool for Embedding-Based Measurement of Data Diversity

Cantao Su, Menan Velayuthan, Esther Ploeger +2

There is growing evidence that data diversity is crucial for developing fair and robust NLP models. However, current approaches to measure diversity remain inconsistent and fragmen…

cs.LG2026

Confusion-Aware Transfer Teacher Curriculum Learning Framework: Disentangling Scoring and Pacing Effects

Savini Kommalage, Sanka Mohottala, Asiri Gawesha +5

Curriculum learning couples two design choices, how samples are scored by difficulty and how harder samples are paced into training, making it difficult to attribute observed gains…

cs.CL2025

From Phonemes to Meaning: Evaluating Large Language Models on Tamil

Jeyarajalingam Varsha, Menan Velayuthan, Sumirtha Karunakaran +2

Large Language Models (LLMs) have shown strong generalization across tasks in high-resource languages; however, their linguistic competence in low-resource and morphologically rich…

cs.CL2025

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…

cs.CL2024

Egalitarian Language Representation in Language Models: It All Begins with Tokenizers

Menan Velayuthan, Kengatharaiyer Sarveswaran

Tokenizers act as a bridge between human language and the latent space of language models, influencing how language is represented in these models. Due to the immense popularity of…