1 citations · 1 across the 3 of their papers we have counts for
3 papers · 1 filter
Comparative analysis of subword tokenization approaches for Indian languages
Sudhansu Bala Das, Samujjal Choudhury, Tapas Kumar Mishra +1
Tokenization is the act of breaking down text into smaller parts, or tokens, that are easier for machines to process. This is a key phase in machine translation (MT) models. Subwor…
An approach for mistranslation removal from popular dataset for Indic MT Task
Sudhansu Bala Das, Leo Raphael Rodrigues, Tapas Kumar Mishra +1
The conversion of content from one language to another utilizing a computer system is known as Machine Translation (MT). Various techniques have come up to ensure effective transla…
Part-of-Speech Tagging of Odia Language Using statistical and Deep Learning-Based Approaches
Tusarkanta Dalai, Tapas Kumar Mishra, Pankaj K Sa
Automatic Part-of-speech (POS) tagging is a preprocessing step of many natural language processing (NLP) tasks such as name entity recognition (NER), speech processing, information…