Unsupervised detection of coordinated information operations in the wild
arXiv:2401.06205 · doi:10.1140/epjds/s13688-025-00544-y
Abstract
This paper introduces and tests an unsupervised method for detecting novel coordinated inauthentic information operations (CIOs) in realistic settings. This method uses Bayesian inference to identify groups of accounts that share similar account-level characteristics and target similar narratives. We solve the inferential problem using amortized variational inference, allowing us to efficiently infer group identities for millions of accounts. We validate this method using a set of five CIOs from three countries discussing four topics on Twitter. Our unsupervised approach increases detection power (area under the precision-recall curve) relative to a naive baseline (by a factor of 76 to 580), relative to the use of simple flags or narratives on their own (by a factor of 1.3 to 4.8), and comes quite close to a supervised benchmark. Our method is robust to observing only a small share of messaging on the topic, having only weak markers of inauthenticity, and to the CIO accounts making up a tiny share of messages and accounts on the topic. Although we evaluate the results on Twitter, the method is general enough to be applied in many social-media settings.
34 pages, 10 figures
References in corpus (5)
- Unveiling Coordinated Groups Behind White Helmets Disinformation
- Gender Classification and Bias Mitigation in Facial Images
- An Automated Pipeline for Character and Relationship Extraction from Readers' Literary Book Reviews on Goodreads.com
- Temporal Dynamics of Coordinated Online Behavior: Stability, Archetypes, and Influence
- Exposing Influence Campaigns in the Age of LLMs: A Behavioral-Based AI Approach to Detecting State-Sponsored Trolls