An Army of Me: Sockpuppets in Online Discussion Communities
arXiv:1703.07355 · doi:10.1145/3038912.3052677
Abstract
In online discussion communities, users can interact and share information and opinions on a wide variety of topics. However, some users may create multiple identities, or sockpuppets, and engage in undesired behavior by deceiving others or manipulating discussions. In this work, we study sockpuppetry across nine discussion communities, and show that sockpuppets differ from ordinary users in terms of their posting behavior, linguistic traits, as well as social network structure. Sockpuppets tend to start fewer discussions, write shorter posts, use more personal pronouns such as "I", and have more clustered ego-networks. Further, pairs of sockpuppets controlled by the same individual are more likely to interact on the same discussion at the same time than pairs of ordinary users. Our analysis suggests a taxonomy of deceptive behavior in discussion communities. Pairs of sockpuppets can vary in their deceptiveness, i.e., whether they pretend to be different users, or their supportiveness, i.e., if they support arguments of other sockpuppets controlled by the same user. We apply these findings to a series of prediction tasks, notably, to identify whether a pair of accounts belongs to the same underlying user or not. Altogether, this work presents a data-driven view of deception in online discussion communities and paves the way towards the automatic detection of sockpuppets.
26th International World Wide Web conference 2017 (WWW 2017)
References in corpus (1)
Cited by in corpus (17)
- Data Governance in the Age of Large-Scale Data-Driven Language Technology
- Towards offensive language detection and reduction in four Software Engineering communities
- User Identity Linkage in Social Media Using Linguistic and Social Interaction Features
- Analysing user identity via time-sensitive semantic edit distance (t-SED): A case study of Russian trolls on Twitter
- Characterizing, Detecting, and Predicting Online Ban Evasion
- PETGEN: Personalized Text Generation Attack on Deep Sequence Embedding-based Classification Models
- Wikipedia in Wartime: Experiences of Wikipedians Maintaining Articles About the Russia-Ukraine War
- A preliminary approach to knowledge integrity risk assessment in Wikipedia projects
- TrollsWithOpinion: A Dataset for Predicting Domain-specific Opinion Manipulation in Troll Memes
- Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information Operations
- Co-Membership-based Generic Anomalous Communities Detection
- From Symbols to Embeddings: A Tale of Two Representations in Computational Social Science
- EmoMed: An Emotionally-Aware Agent for Multimodal Medical Support with Real-Time Information Retrieval
- Analysing Russian Trolls via NLP tools
- Improving Authorship Verification using Linguistic Divergence
- Deep Unified Multimodal Embeddings for Understanding both Content and Users in Social Media Networks
- Towards Trustworthy Deception Detection: Benchmarking Model Robustness across Domains, Modalities, and Languages