3 papers
cs.CL2020
4chan & 8chan embeddings
Pierre Voué, Tom De Smedt, Guy De Pauw
We have collected over 30M messages from the publicly available /pol/ message boards on 4chan and 8chan, and compiled them into a model of toxic language use. The trained word embe…
cs.CL2018
Multilingual Cross-domain Perspectives on Online Hate Speech
Tom De Smedt, Sylvia Jaki, Eduan Kotzé +4
In this report, we present a study of eight corpora of online hate speech, by demonstrating the NLP techniques that we used to collect and analyze the jihadist, extremist, racist,…
cs.CL2018
Automatic Detection of Online Jihadist Hate Speech
Tom De Smedt, Guy De Pauw, Pieter Van Ostaeyen
We have developed a system that automatically detects online jihadist hate speech with over 80% accuracy, by using techniques from Natural Language Processing and Machine Learning.…