Fame for sale: efficient detection of fake Twitter followers
arXiv:1509.04098 · doi:10.1016/j.dss.2015.09.003
Abstract
are those Twitter accounts specifically created to inflate the number of followers of a target account. Fake followers are dangerous for the social platform and beyond, since they may alter concepts like popularity and influence in the Twittersphere - hence impacting on economy, politics, and society. In this paper, we contribute along different dimensions. First, we review some of the most relevant existing features and rules (proposed by Academia and Media) for anomalous Twitter accounts detection. Second, we create a baseline dataset of verified human and fake follower accounts. Such baseline dataset is publicly available to the scientific community. Then, we exploit the baseline dataset to train a set of machine-learning classifiers built over the reviewed rules and features. Our results show that most of the rules proposed by Media provide unsatisfactory performance in revealing fake followers, while features proposed in the past by Academia for spam detection provide good results. Building on the most promising features, we revise the classifiers both in terms of reduction of overfitting and cost for gathering the data needed to compute the features. The final result is a novel classifier, general enough to thwart overfitting, lightweight thanks to the usage of the less costly features, and still able to correctly classify more than 95% of the accounts of the original training set. We ultimately perform an information fusion-based sensitivity analysis, to assess the global sensitivity of each of the features employed by the classifier. The findings reported in this paper, other than being supported by a thorough experimental methodology and interesting on their own, also pave the way for further investigation on the novel issue of fake Twitter followers.
References in corpus (2)
Cited by in corpus (40)
- Arming the public with artificial intelligence to counter social bots
- The paradigm-shift of social spambots: Evidence, theories, and tools for the arms race
- A Decade of Social Bot Detection
- TwiBot-20: A Comprehensive Twitter Bot Detection Benchmark
- DNA-inspired online behavioral modeling and its application to spambot detection
- Social Fingerprinting: detection of spambot groups through DNA-inspired behavioral modeling
- Detection of Novel Social Bots by Ensembles of Specialized Classifiers
- Instagram Fake and Automated Account Detection
- The role of bot squads in the political propaganda on Twitter
- Cashtag piggybacking: uncovering spam and bot activity in stock microblogs on Twitter
- SATAR: A Self-supervised Approach to Twitter Account Representation Learning and its Application in Bot Detection
- Unsupervised Social Bot Detection via Structural Information Theory
- RoSGAS: Adaptive Social Bot Detection with Reinforced Self-Supervised GNN Architecture Search
- Flow of online misinformation during the peak of the COVID-19 pandemic in Italy
- Influence of augmented humans in online interactions during voting events
- The Looming Threat of Fake and LLM-generated LinkedIn Profiles: Challenges and Opportunities for Detection and Prevention
- Analysing Twitter Semantic Networks: the case of 2018 Italian Elections
- Brexit and bots: characterizing the behaviour of automated accounts on Twitter during the UK election
- Simplistic Collection and Labeling Practices Limit the Utility of Benchmark Datasets for Twitter Bot Detection
- On the efficacy of old features for the detection of new bots
- Sustainable Development Goals as unifying narratives in large UK firms' Twitter discussions
- Bow-Tie Structures of Twitter Discursive Communities
- Fine-Grained Prediction of Political Leaning on Social Media with Unsupervised Deep Learning
- Better Safe Than Sorry: An Adversarial Approach to Improve Social Bot Detection
- A model for the Twitter sentiment curve
- CasCIFF: A Cross-Domain Information Fusion Framework Tailored for Cascade Prediction in Social Networks
- The Anatomy of Conspirators: Unveiling Traits using a Comprehensive Twitter Dataset
- MediaRank: Computational Ranking of Online News Sources
- Analyzing Activity and Suspension Patterns of Twitter Bots Attacking Turkish Twitter Trends by a Longitudinal Dataset
- How to Use Graph Data in the Wild to Help Graph Anomaly Detection?
- Beyond Trial-and-Error: Predicting User Abandonment After a Moderation Intervention
- How many bots are you following?
- Relevance-Aware Anomalous Users Detection in Social Network via Graph Neural Network
- Ensuring the Inclusive Use of Natural Language Processing in the Global Response to COVID-19
- Six Million (Suspected) Fake Stars in GitHub: A Growing Spiral of Popularity Contests, Spams, and Malware
- Forecasting Political News Engagement on Social Media
- Analyzing time series activity of Twitter political spambots
- Italian Twitter semantic network during the Covid-19 epidemic
- Characterizing Online Criticism of Partisan News Media Using Weakly Supervised Learning
- Prioritizing Original News on Facebook