1 citations · 1 across the 2 of their papers we have counts for
3 papers · 1 filter
Comparative analysis of subword tokenization approaches for Indian languages
Sudhansu Bala Das, Samujjal Choudhury, Tapas Kumar Mishra +1
Tokenization is the act of breaking down text into smaller parts, or tokens, that are easier for machines to process. This is a key phase in machine translation (MT) models. Subwor…
An approach for mistranslation removal from popular dataset for Indic MT Task
Sudhansu Bala Das, Leo Raphael Rodrigues, Tapas Kumar Mishra +1
The conversion of content from one language to another utilizing a computer system is known as Machine Translation (MT). Various techniques have come up to ensure effective transla…
Multilingual Neural Machine Translation System for Indic to Indic Languages
Sudhansu Bala Das, Divyajyoti Panda, Tapas Kumar Mishra +2
This paper gives an Indic-to-Indic (IL-IL) MNMT baseline model for 11 ILs implemented on the Samanantar corpus and analyzed on the Flores-200 corpus. All the models are evaluated u…