1 citations · 1 across the 5 of their papers we have counts for
6 papers
Typologically-Informed Candidate Reranking for LLM-based Translation into Low-Resource Languages
Nipuna Abeykoon, Ashen Weerathunga, Pubudu Wijesinghe +1
Large language models trained predominantly on high-resource languages exhibit systematic biases toward dominant typological patterns, leading to structural non-conformance when tr…
Decoding Fake Narratives in Spreading Hateful Stories: A Dual-Head RoBERTa Model with Multi-Task Learning
Yash Bhaskar, Sankalp Bahad, Parameswari Krishnamurthy
Social media platforms, while enabling global connectivity, have become hubs for the rapid spread of harmful content, including hate speech and fake narratives \cite{davidson2017au…
Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
Yash Bhaskar, Parameswari Krishnamurthy
This paper presents the systems submitted by the Yes-MT team for the Low-Resource Indic Language Translation Shared Task at WMT 2024 (Pakray et al., 2024), focusing on translating…
CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems
Soham Bhattacharjee, Mukund K Roy, Yathish Poojary +19
India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognize…
Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
Saketh Reddy Vemula, Sandipan Dandapat, Dipti Misra Sharma +1
The relationship between tokenizer algorithm (e.g., Byte-Pair Encoding (BPE), Unigram), morphological alignment, tokenization quality (e.g., compression efficiency), and downstream…
keepitsimple at SemEval-2025 Task 3: LLM-Uncertainty based Approach for Multilingual Hallucination Span Detection
Saketh Reddy Vemula, Parameswari Krishnamurthy
Identification of hallucination spans in black-box language model generated text is essential for applications in the real world. A recent attempt at this direction is SemEval-2025…