most citedYes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024

1 citations · 1 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CL2026

Typologically-Informed Candidate Reranking for LLM-based Translation into Low-Resource Languages

Nipuna Abeykoon, Ashen Weerathunga, Pubudu Wijesinghe +1

Large language models trained predominantly on high-resource languages exhibit systematic biases toward dominant typological patterns, leading to structural non-conformance when tr…

cs.CL2025

Decoding Fake Narratives in Spreading Hateful Stories: A Dual-Head RoBERTa Model with Multi-Task Learning

Yash Bhaskar, Sankalp Bahad, Parameswari Krishnamurthy

Social media platforms, while enabling global connectivity, have become hubs for the rapid spread of harmful content, including hate speech and fake narratives \cite{davidson2017au…

cs.CL20251 cited

Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024

Yash Bhaskar, Parameswari Krishnamurthy

This paper presents the systems submitted by the Yes-MT team for the Low-Resource Indic Language Translation Shared Task at WMT 2024 (Pakray et al., 2024), focusing on translating…

cs.CL2025

CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems

Soham Bhattacharjee, Mukund K Roy, Yathish Poojary +19

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognize…

cs.CL2025

Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment

Saketh Reddy Vemula, Sandipan Dandapat, Dipti Misra Sharma +1

The relationship between tokenizer algorithm (e.g., Byte-Pair Encoding (BPE), Unigram), morphological alignment, tokenization quality (e.g., compression efficiency), and downstream…

cs.CL2025

keepitsimple at SemEval-2025 Task 3: LLM-Uncertainty based Approach for Multilingual Hallucination Span Detection

Saketh Reddy Vemula, Parameswari Krishnamurthy

Identification of hallucination spans in black-box language model generated text is essential for applications in the real world. A recent attempt at this direction is SemEval-2025…