2 papers
cs.CL2026
Mitrasamgraha: A Comprehensive Classical Sanskrit Machine Translation Dataset
Sebastian Nehrdich, David Allport, Sven Sellmer +5
While machine translation is regarded as a "solved problem" for many high-resource languages, close analysis quickly reveals that this is not the case for content that shows challe…
cs.CL2026
MITRA: A Large-Scale Parallel Corpus and Multilingual Pretrained Language Model for Machine Translation and Semantic Retrieval for PÄli, Sanskrit, Buddhist Chinese, and Tibetan
Sebastian Nehrdich, Kurt Keutzer
Ancient Buddhist literature features frequent, yet often unannotated, textual parallels spread across diverse languages: Sanskrit, PÄli, Buddhist Chinese, Tibetan, and more. The s…