An Effective, Robust and Fairness-aware Hate Speech Detection Framework
arXiv:2409.17191 · doi:10.1109/BigData52589.2021.9672022
Abstract
With the widespread online social networks, hate speeches are spreading faster and causing more damage than ever before. Existing hate speech detection methods have limitations in several aspects, such as handling data insufficiency, estimating model uncertainty, improving robustness against malicious attacks, and handling unintended bias (i.e., fairness). There is an urgent need for accurate, robust, and fair hate speech classification in online social networks. To bridge the gap, we design a data-augmented, fairness addressed, and uncertainty estimated novel framework. As parts of the framework, we propose Bidirectional Quaternion-Quasi-LSTM layers to balance effectiveness and efficiency. To build a generalized model, we combine five datasets collected from three platforms. Experiment results show that our model outperforms eight state-of-the-art methods under both no attack scenario and various attack scenarios, indicating the effectiveness and robustness of our model. We share our code along with combined dataset for better future research
IEEE BigData 2021
References in corpus (19)
- Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning
- Equality of Opportunity in Supervised Learning
- Unsupervised Data Augmentation for Consistency Training
- The Curious Case of Neural Text Degeneration
- Deep Learning for Hate Speech Detection in Tweets
- TextBugger: Generating Adversarial Text Against Real-world Applications
- Quasi-Recurrent Neural Networks
- Data Decisions and Theoretical Implications when Adversarially Learning Fair Representations
- A General Framework for Uncertainty Estimation in Deep Learning
- Ex Machina: Personal Attacks Seen at Scale
- QuaterNet: A Quaternion-based Recurrent Model for Human Motion
- Modeling Human Motion with Quaternion-based Neural Networks
- Stereotypical Bias Removal for Hate Speech Detection Task using Knowledge-based Generalizations
- A Simple but Tough-to-Beat Data Augmentation Approach for Natural Language Understanding and Generation
- CONAN -- COunter NArratives through Nichesourcing: a Multilingual Dataset of Responses to Fight Online Hate Speech
- Adv-BERT: BERT is not robust on misspellings! Generating nature adversarial samples on BERT
- Evaluation of Neural Architectures Trained with Square Loss vs Cross-Entropy in Classification Tasks
- Quaternion-Based Self-Attentive Long Short-Term User Preference Encoding for Recommendation
- SWE2: SubWord Enriched and Significant Word Emphasized Framework for Hate Speech Detection