activity
20202024
most citedDomain-specific MT for Low-resource Languages: The case of Bambara-French

4 citations · 9 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2024

ARTICLE: Annotator Reliability Through In-Context Learning

Sujan Dutta, Deepak Pandita, Tharindu Cyril Weerasooriya +3

Ensuring annotator quality in training and evaluation data is a key piece of machine learning in NLP. Tasks such as sentiment analysis and offensive speech detection are intrinsica…

cs.CL20214 cited

Cross-lingual Offensive Language Identification for Low Resource Languages: The Case of Marathi

Saurabh Gaikwad, Tharindu Ranasinghe, Marcos Zampieri +1

The widespread presence of offensive language on social media motivated the development of systems capable of recognizing such content automatically. Apart from a few notable excep…

cs.CL2021

LCP-RIT at SemEval-2021 Task 1: Exploring Linguistic Features for Lexical Complexity Prediction

Abhinandan Desai, Kai North, Marcos Zampieri +1

This paper describes team LCP-RIT's submission to the SemEval-2021 Task 1: Lexical Complexity Prediction (LCP). The task organizers provided participants with an augmented version…

cs.CL20214 cited

Domain-specific MT for Low-resource Languages: The case of Bambara-French

Allahsera Auguste Tapo, Michael Leventhal, Sarah Luger +2

Translating to and from low-resource languages is a challenge for machine translation (MT) systems due to a lack of parallel data. In this paper we address the issue of domain-spec…

cs.CL2020

Assessing Human Translations from French to Bambara for Machine Learning: a Pilot Study

Michael Leventhal, Allahsera Tapo, Sarah Luger +2

We present novel methods for assessing the quality of human-translated aligned texts for learning machine translation models of under-resourced languages. Malian university student…