most citedA simple language-agnostic yet very strong baseline system for hate speech and offensive content identification

1 citations · 1 across the 5 of their papers we have counts for

collaborators

7 papers

cs.CL2022

Please, Don't Forget the Difference and the Confidence Interval when Seeking for the State-of-the-Art Status

Yves Bestgen

This paper argues for the widest possible use of bootstrap confidence intervals for comparing NLP system performances instead of the state-of-the-art status (SOTA) and statistical…

cs.CL2022

SATLab at SemEval-2022 Task 4: Trying to Detect Patronizing and Condescending Language with only Character and Word N-grams

Yves Bestgen

A logistic regression model only fed with character and word n-grams is proposed for the SemEval-2022 Task 4 on Patronizing and Condescending Language Detection (PCL). It obtained…

cs.CL20221 cited

A simple language-agnostic yet very strong baseline system for hate speech and offensive content identification

Yves Bestgen

For automatically identifying hate speech and offensive content in tweets, a system based on a classical supervised algorithm only fed with character n-grams, and thus completely l…

cs.CL2021

Using CollGram to Compare Formulaic Language in Human and Neural Machine Translation

Yves Bestgen

A comparison of formulaic sequences in human and neural machine translation of quality newspaper articles shows that neural machine translations contain less lower-frequency, but s…

cs.CL2021

LAST at SemEval-2021 Task 1: Improving Multi-Word Complexity Prediction Using Bigram Association Measures

Yves Bestgen

This paper describes the system developed by the Laboratoire d'analyse statistique des textes (LAST) for the Lexical Complexity Prediction shared task at SemEval-2021. The proposed…

cs.CL2021

Using Fisher's Exact Test to Evaluate Association Measures for N-grams

Yves Bestgen

To determine whether some often-used lexical association measures assign high scores to n-grams that chance could have produced as frequently as observed, we used an extension of F…