most citedDomain-specific MT for Low-resource Languages: The case of Bambara-French

4 citations · 9 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL20214 cited

Cross-lingual Offensive Language Identification for Low Resource Languages: The Case of Marathi

Saurabh Gaikwad, Tharindu Ranasinghe, Marcos Zampieri +1

The widespread presence of offensive language on social media motivated the development of systems capable of recognizing such content automatically. Apart from a few notable excep…

cs.AI20211 cited

Improving Label Quality by Jointly Modeling Items and Annotators

Tharindu Cyril Weerasooriya, Alexander G. Ororbia, Christopher M. Homan

We propose a fully Bayesian framework for learning ground truth labels from noisy annotators. Our framework ensures scalability by factoring a generative, Bayesian soft clustering…

cs.CL2021

LCP-RIT at SemEval-2021 Task 1: Exploring Linguistic Features for Lexical Complexity Prediction

Abhinandan Desai, Kai North, Marcos Zampieri +1

This paper describes team LCP-RIT's submission to the SemEval-2021 Task 1: Lexical Complexity Prediction (LCP). The task organizers provided participants with an augmented version…

cs.CL20214 cited

Domain-specific MT for Low-resource Languages: The case of Bambara-French

Allahsera Auguste Tapo, Michael Leventhal, Sarah Luger +2

Translating to and from low-resource languages is a challenge for machine translation (MT) systems due to a lack of parallel data. In this paper we address the issue of domain-spec…

cs.CL2020

Assessing Human Translations from French to Bambara for Machine Learning: a Pilot Study

Michael Leventhal, Allahsera Tapo, Sarah Luger +2

We present novel methods for assessing the quality of human-translated aligned texts for learning machine translation models of under-resourced languages. Malian university student…