activity
20172022
most citedExploring the Use of Text Classification in the Legal Domain

105 citations · 261 across the 25 of their papers we have counts for

collaborators

47 papers

cs.CL2022

Predicting the Type and Target of Offensive Social Media Posts in Marathi

Marcos Zampieri, Tharindu Ranasinghe, Mrinal Chaudhari +4

The presence of offensive language on social media is very common motivating platforms to invest in strategies to make communities safer. This includes developing robust machine le…

cs.CL20222 cited

Overview of the HASOC Subtrack at FIRE 2022: Offensive Language Identification in Marathi

Tharindu Ranasinghe, Kai North, Damith Premasiri +1

The widespread of offensive content online has become a reason for great concern in recent years, motivating researchers to develop robust systems capable of identifying such conte…

cs.CL2022

Lexical Simplification Benchmarks for English, Portuguese, and Spanish

Sanja Stajner, Daniel Ferres, Matthew Shardlow +3

Even in highly-developed countries, as many as 15-30\% of the population can only understand texts written using a basic vocabulary. Their understanding of everyday texts is limite…

cs.CL2021

FBERT: A Neural Transformer for Identifying Offensive Content

Diptanu Sarkar, Marcos Zampieri, Tharindu Ranasinghe +1

Transformer-based models such as BERT, XLNET, and XLM-R have achieved state-of-the-art performance across various NLP tasks including the identification of offensive language and h…

cs.CL20214 cited

Cross-lingual Offensive Language Identification for Low Resource Languages: The Case of Marathi

Saurabh Gaikwad, Tharindu Ranasinghe, Marcos Zampieri +1

The widespread presence of offensive language on social media motivated the development of systems capable of recognizing such content automatically. Apart from a few notable excep…

cs.SE202118 cited

An Ensemble Approach for Annotating Source Code Identifiers with Part-of-speech Tags

Christian D. Newman, Michael J. Decker, Reem S. AlSuhaibani +7

This paper presents an ensemble part-of-speech tagging approach for source code identifiers. Ensemble tagging is a technique that uses machine-learning and the output from multiple…