3 papers
cs.CL2018
Free as in Free Word Order: An Energy Based Model for Word Segmentation and Morphological Tagging in Sanskrit
Amrith Krishna, Bishal Santra, Sasi Prasanth Bandaru +4
The configurational information in sentences of a free word order language such as Sanskrit is of limited use. Thus, the context of the entire sentence will be desirable even for b…
cs.CL2018
Upcycle Your OCR: Reusing OCRs for Post-OCR Text Correction in Romanised Sanskrit
Amrith Krishna, Bodhisattwa Prasad Majumder, Rajesh Shreedhar Bhat +1
We propose a post-OCR text correction approach for digitising texts in Romanised Sanskrit. Owing to the lack of resources our approach uses OCR models trained for other languages w…
cs.CL2018
Building a Word Segmenter for Sanskrit Overnight
Vikas Reddy, Amrith Krishna, Vishnu Dutt Sharma +3
There is an abundance of digitised texts available in Sanskrit. However, the word segmentation task in such texts are challenging due to the issue of 'Sandhi'. In Sandhi, words in…