From the 1 of 10 linked papers with an AI index.
10 papers
Language Models are not Equally Robust to Non-Canonical Tokenization across Languages
Poulami Ghosh, Preethi Jyothi
The paper investigates how language models handle non-canonical tokenizations across 27 languages, finding that robustness varies by language and token fragmentation, and shows tha…
Are Multilingual Models Actually Improving? Isolating True Cross-Lingual Transfer
Prasoon Bajpai, Eleftheria Briakou, Colin Cherry +2
Cross-lingual transfer is a model's ability to generalize capabilities from well-represented source languages to under-represented target languages. Existing measures of a model's…
XITE: Cross-lingual Interpolation for Transfer using Embeddings
Barah Fazili, Preethi Jyothi
Facilitating cross-lingual transfer in multilingual language models remains a critical challenge. Towards this goal, we propose an embedding-based data augmentation technique calle…
Evaluating Extremely Low-Resource Machine Translation: A Comparative Study of ChrF++ and BLEU Metrics
Sanjeev Kumar, Preethi Jyothi, Pushpak Bhattacharyya
Evaluating machine translation (MT) quality in extremely low-resource language (ELRL) scenarios poses unique challenges, as widely used metrics such as BLEU, effective in high-reso…
LASER: An LLM-based ASR Scoring and Evaluation Rubric
Amruta Parulekar, Preethi Jyothi
Standard ASR evaluation metrics like Word Error Rate (WER) tend to unfairly penalize morphological and syntactic nuances that do not significantly alter sentence semantics. We intr…
LEVOS: Leveraging Vocabulary Overlap with Sanskrit to Generate Technical Lexicons in Indian Languages
Karthika N J, Krishnakant Bhatt, Ganesh Ramakrishnan +1
Translating technical terms into lexically similar, low-resource Indian languages remains a challenge due to limited parallel data and the complexity of linguistic structures. We p…