4 citations · 9 across the 4 of their papers we have counts for
5 papers
Cross-lingual Offensive Language Identification for Low Resource Languages: The Case of Marathi
Saurabh Gaikwad, Tharindu Ranasinghe, Marcos Zampieri +1
The widespread presence of offensive language on social media motivated the development of systems capable of recognizing such content automatically. Apart from a few notable excep…
Improving Label Quality by Jointly Modeling Items and Annotators
Tharindu Cyril Weerasooriya, Alexander G. Ororbia, Christopher M. Homan
We propose a fully Bayesian framework for learning ground truth labels from noisy annotators. Our framework ensures scalability by factoring a generative, Bayesian soft clustering…
LCP-RIT at SemEval-2021 Task 1: Exploring Linguistic Features for Lexical Complexity Prediction
Abhinandan Desai, Kai North, Marcos Zampieri +1
This paper describes team LCP-RIT's submission to the SemEval-2021 Task 1: Lexical Complexity Prediction (LCP). The task organizers provided participants with an augmented version…
Domain-specific MT for Low-resource Languages: The case of Bambara-French
Allahsera Auguste Tapo, Michael Leventhal, Sarah Luger +2
Translating to and from low-resource languages is a challenge for machine translation (MT) systems due to a lack of parallel data. In this paper we address the issue of domain-spec…
Assessing Human Translations from French to Bambara for Machine Learning: a Pilot Study
Michael Leventhal, Allahsera Tapo, Sarah Luger +2
We present novel methods for assessing the quality of human-translated aligned texts for learning machine translation models of under-resourced languages. Malian university student…