Exploring limits to prediction in complex social systems
arXiv:1602.01013 · doi:10.1145/2872427.2883001
Abstract
How predictable is success in complex social systems? In spite of a recent profusion of prediction studies that exploit online social and information network data, this question remains unanswered, in part because it has not been adequately specified. In this paper we attempt to clarify the question by presenting a simple stylized model of success that attributes prediction error to one of two generic sources: insufficiency of available data and/or models on the one hand; and inherent unpredictability of complex social systems on the other. We then use this model to motivate an illustrative empirical study of information cascade size prediction on Twitter. Despite an unprecedented volume of information about users, content, and past performance, our best performing models can explain less than half of the variance in cascade sizes. In turn, this result suggests that even with unlimited data predictive performance would be bounded well below deterministic accuracy. Finally, we explore this potential bound theoretically using simulations of a diffusion process on a random scale free network similar to Twitter. We show that although higher predictive power is possible in theory, such performance requires a homogeneous system and perfect ex-ante knowledge of it: even a small degree of uncertainty in estimating product quality or slight variation in quality across products leads to substantially more restrictive bounds on predictability. We conclude that realistic bounds on predictive accuracy are not dissimilar from those we have obtained empirically, and that such bounds for other complex social systems for which data is more difficult to obtain are likely even lower.
12 pages, 6 figures, Proceedings of the 25th ACM International World Wide Web Conference (WWW) 2016
References in corpus (5)
Cited by in corpus (26)
- A Survey of Information Cascade Analysis: Models, Predictions, and Recent Advances
- Feature Driven and Point Process Approaches for Popularity Prediction
- Expecting to be HIP: Hawkes Intensity Processes for Social Media Popularity
- SIR-Hawkes: Linking Epidemic Models and Hawkes Processes to Model Diffusions in Finite Populations
- Influence of augmented humans in online interactions during voting events
- Luck is Hard to Beat: The Difficulty of Sports Prediction
- Predicting and Understanding Law-Making with Word Vectors and an Ensemble Model
- Beyond network centrality: Individual-level behavioral traits for predicting information superspreaders in social media
- Will This Video Go Viral? Explaining and Predicting the Popularity of Youtube Videos
- Sequential Prediction of Social Media Popularity with Deep Temporal Context Networks
- Postmortem memory of public figures in news and social media
- Modeling Popularity in Asynchronous Social Media Streams with Recurrent Neural Networks
- Popularity Prediction on Social Platforms with Coupled Graph Neural Networks
- More than Meets the Tie: Examining the Role of Interpersonal Relationships in Social Networks
- Anticipated versus Actual Effects of Platform Design Change: A Case Study of Twitter's Character Limit
- Variation across Scales: Measurement Fidelity under Twitter Data Sampling
- Modeling Information Cascades with Self-exciting Processes via Generalized Epidemic Models
- SMP Challenge: An Overview of Social Media Prediction Challenge 2019
- The wisdom of the few: Predicting collective success from individual behavior
- Structural patterns of information cascades and their implications for dynamics and semantics
- An influencer-based approach to understanding radical right viral tweets
- Learning Representations of Social Media Users
- Utilizing Citation Network Structure to Predict Citation Counts: A Deep Learning Approach
- Estimating Attention Flow in Online Video Networks
- Stochastic differential theory of cricket
- Common Growth Patterns for Regional Social Networks: a Point Process Approach