9 papers
What Tokens are Learned when Tokenization is Optimized Jointly with Language Modeling?
Saketh Reddy Vemula, Parameswari Krishnamurthy
Tokenization is a fundamental component of language modeling pipelines. Despite its importance, it is often fixed, even though it significantly impacts model performance across lan…
SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs
Debopriyo Banerjee, Kapil Rajesh Kavitha, Angana Borah +11
Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally…
Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages
Saeed Almheiri, Bilal Elbouardi, Salsabila Zahirah Pranida +16
Idiomatic expressions pose a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpreta…
Typologically-Informed Candidate Reranking for LLM-based Translation into Low-Resource Languages
Nipuna Abeykoon, Ashen Weerathunga, Pubudu Wijesinghe +1
Large language models trained predominantly on high-resource languages exhibit systematic biases toward dominant typological patterns, leading to structural non-conformance when tr…
Decoding Fake Narratives in Spreading Hateful Stories: A Dual-Head RoBERTa Model with Multi-Task Learning
Yash Bhaskar, Sankalp Bahad, Parameswari Krishnamurthy
Social media platforms, while enabling global connectivity, have become hubs for the rapid spread of harmful content, including hate speech and fake narratives \cite{davidson2017au…
Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
Yash Bhaskar, Parameswari Krishnamurthy
This paper presents the systems submitted by the Yes-MT team for the Low-Resource Indic Language Translation Shared Task at WMT 2024 (Pakray et al., 2024), focusing on translating…