Echoes of Norms: Investigating Counterspeech Bots' Influence on Bystanders in Online Communities
arXiv:2603.03687 · doi:10.1145/3772318.3791797
Abstract
Counterspeech offers a non-repressive approach to moderate hate speech in online communities. Research has examined how counterspeech chatbots restrain hate speakers and support targets, but their impact on bystanders remains unclear. Therefore, we developed a counterspeech strategy framework and built \textit{Civilbot} for a mixed-method within-subjects study. Bystanders generally viewed Civilbot as credible and normative, though its shallow reasoning limited persuasiveness. Its behavioural effects were subtle: when performing well, it could guide participation or act as a stand-in; when performing poorly, it could discourage bystanders or motivate them to step in. Strategy proved critical: cognitive strategies that appeal to reason, especially when paired with a positive tone, were relatively effective, while mismatch of contexts and strategies could weaken impact. Based on these findings, we offer design insights for mobilizing bystanders and shaping online discourse, highlighting when to intervene and how to do so through reasoning-driven and context-aware strategies.
Accepted to the CHI Conference on Human Factors in Computing Systems (CHI 2026)
References in corpus (21)
- Anyone Can Become a Troll: Causes of Trolling Behavior in Online Discussions
- The Structure of Toxic Conversations on Twitter
- CONAN -- COunter NArratives through Nichesourcing: a Multilingual Dataset of Responses to Fight Online Hate Speech
- "You have to prove the threat is real": Understanding the needs of Female Journalists and Activists to Document and Report Online Harassment
- Proactive Moderation of Online Discussions: Existing Practices and the Potential for Algorithmic Support
- Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
- DeMod: A Holistic Tool with Explainable Detection and Personalized Modification for Toxicity Censorship
- Generate, Prune, Select: A Pipeline for Counterspeech Generation against Online Hate Speech
- Detecting East Asian Prejudice on Social Media
- Perceiving and Countering Hate: The Role of Identity in Online Responses
- Contextualized Counterspeech: Strategies for Adaptation, Personalization, and Evaluation
- Sadness, Anger, or Anxiety: Twitter Users' Emotional Responses to Toxicity in Public Conversations
- Generative AI may backfire for counterspeech
- Outcome-Constrained Large Language Models for Countering Hate Speech
- Designing Human-AI Collaboration to Support Learning in Counterspeech Writing
- Adaptive Human-Agent Teaming: A Review of Empirical Studies from the Process Dynamics Perspective
- Hate Speech and Counter Speech Detection: Conversational Context Does Matter
- Understanding Counterspeech for Online Harm Mitigation
- Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues
- CrowdCounter: A benchmark type-specific multi-target counterspeech dataset
- PANDA -- Paired Anti-hate Narratives Dataset from Asia: Using an LLM-as-a-Judge to Create the First Chinese Counterspeech Dataset