works on

From the 1 of 10 linked papers with an AI index.

collaborators

10 papers

cs.CL2026

Language Models are not Equally Robust to Non-Canonical Tokenization across Languages

Poulami Ghosh, Preethi Jyothi

The paper investigates how language models handle non-canonical tokenizations across 27 languages, finding that robustness varies by language and token fragmentation, and shows tha…

cs.CL2026

Are Multilingual Models Actually Improving? Isolating True Cross-Lingual Transfer

Prasoon Bajpai, Eleftheria Briakou, Colin Cherry +2

Cross-lingual transfer is a model's ability to generalize capabilities from well-represented source languages to under-represented target languages. Existing measures of a model's…

cs.CL2026

XITE: Cross-lingual Interpolation for Transfer using Embeddings

Barah Fazili, Preethi Jyothi

Facilitating cross-lingual transfer in multilingual language models remains a critical challenge. Towards this goal, we propose an embedding-based data augmentation technique calle…

cs.CL2026

Evaluating Extremely Low-Resource Machine Translation: A Comparative Study of ChrF++ and BLEU Metrics

Sanjeev Kumar, Preethi Jyothi, Pushpak Bhattacharyya

Evaluating machine translation (MT) quality in extremely low-resource language (ELRL) scenarios poses unique challenges, as widely used metrics such as BLEU, effective in high-reso…

cs.CL2025

LASER: An LLM-based ASR Scoring and Evaluation Rubric

Amruta Parulekar, Preethi Jyothi

Standard ASR evaluation metrics like Word Error Rate (WER) tend to unfairly penalize morphological and syntactic nuances that do not significantly alter sentence semantics. We intr…

cs.CL2025

LEVOS: Leveraging Vocabulary Overlap with Sanskrit to Generate Technical Lexicons in Indian Languages

Karthika N J, Krishnakant Bhatt, Ganesh Ramakrishnan +1

Translating technical terms into lexically similar, low-resource Indian languages remains a challenge due to limited parallel data and the complexity of linguistic structures. We p…