2 papers
cs.CL2024
MaCmS: Magahi Code-mixed Dataset for Sentiment Analysis
Priya Rani, Gaurav Negi, Theodorus Fransen +1
The present paper introduces new sentiment data, MaCMS, for Magahi-Hindi-English (MHE) code-mixed language, where Magahi is a less-resourced minority language. This dataset is the…
cs.CL2023
Weakly-supervised Deep Cognate Detection Framework for Low-Resourced Languages Using Morphological Knowledge of Closely-Related Languages
Koustava Goswami, Priya Rani, Theodorus Fransen +1
Exploiting cognates for transfer learning in under-resourced languages is an exciting opportunity for language understanding tasks, including unsupervised machine translation, name…