most citedExploring the Use of Text Classification in the Legal Domain

105 citations · 175 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL201733 cited

Detecting Hate Speech in Social Media

Shervin Malmasi, Marcos Zampieri

In this paper we examine methods to detect hate speech in social media, while distinguishing this from general profanity. We aim to establish lexical baselines for this task by app…

cs.CL2017105 cited

Exploring the Use of Text Classification in the Legal Domain

Octavia-Maria Sulea, Marcos Zampieri, Shervin Malmasi +3

In this paper, we investigate the application of text classification methods to support law professionals. We present several experiments applying machine learning techniques to pr…

cs.CL20176 cited

Complex Word Identification: Challenges in Data Annotation and System Performance

Marcos Zampieri, Shervin Malmasi, Gustavo Paetzold +1

This paper revisits the problem of complex word identification (CWI) following up the SemEval CWI shared task. We use ensemble classifiers to investigate how well computational met…

cs.CL2017

Open-Set Language Identification

Shervin Malmasi

We present the first open-set language identification experiments using one-class classification. We first highlight the shortcomings of traditional feature extraction methods and…

cs.CL20175 cited

Including Dialects and Language Varieties in Author Profiling

Alina Maria Ciobanu, Marcos Zampieri, Shervin Malmasi +1

This paper presents a computational approach to author profiling taking gender and language variety into account. We apply an ensemble system with the output of multiple linear SVM…

cs.CL201726 cited

Native Language Identification using Stacked Generalization

Shervin Malmasi, Mark Dras

Ensemble methods using multiple classifiers have proven to be the most successful approach for the task of Native Language Identification (NLI), achieving the current state of the…