ModSandbox: Facilitating Online Community Moderation Through Error Prediction and Improvement of Automated Rules
arXiv:2210.09569 · doi:10.1145/3544548.3581057
Abstract
Despite the common use of rule-based tools for online content moderation, human moderators still spend a lot of time monitoring them to ensure that they work as intended. Based on surveys and interviews with Reddit moderators who use AutoModerator, we identified the main challenges in reducing false positives and false negatives of automated rules: not being able to estimate the actual effect of a rule in advance and having difficulty figuring out how the rules should be updated. To address these issues, we built ModSandbox, a novel virtual sandbox system that detects possible false positives and false negatives of a rule to be improved and visualizes which part of the rule is causing issues. We conducted a user study with online content moderators, finding that ModSandbox can support quickly finding possible false positives and false negatives of automated rules and guide moderators to update those to reduce future errors.
24 pages, 10 figures
References in corpus (3)
Cited by in corpus (4)
- Linguistically Differentiating Acts and Recalls of Racial Microaggressions on Social Media
- DeMod: A Holistic Tool with Explainable Detection and Personalized Modification for Toxicity Censorship
- End User Authoring of Personalized Content Classifiers: Comparing Example Labeling, Rule Writing, and LLM Prompting
- Mapping Community Appeals Systems: Lessons for Community-led Moderation in Multi-Level Governance