4 papers
Naamah: A Large Scale Synthetic Sanskrit NER Corpus via DBpedia Seeding and LLM Generation
Akhil Rajeev P, Annarao Kulkarni
The digitisation of classical Sanskrit literature is impeded by a scarcity of annotated resources, particularly for Named Entity Recognition. While recent methodologies utilise gen…
Chandomitra: Towards Generating Structured Sanskrit Poetry from Natural Language Inputs
Manoj Balaji Jagadeeshan, Samarth Bhatia, Pretam Ray +7
Text Generation has achieved remarkable performance using large language models. It has also been recently well-studied that these large language models are capable of creative gen…
Accent Placement Models for Rigvedic Sanskrit Text
Akhil Rajeev P, Annarao Kulkarni
The Rigveda, among the oldest Indian texts in Vedic Sanskrit, employs a distinctive pitch-accent system : udÄtta, anudÄtta, svarita whose marks encode melodic and interpretive cu…
CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems
Soham Bhattacharjee, Mukund K Roy, Yathish Poojary +19
India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognize…