Early Prediction of Movie Box Office Success based on Wikipedia Activity Big Data
arXiv:1211.0970 · doi:10.1371/journal.pone.0071226
Abstract
Use of socially generated "big data" to access information about collective states of the minds in human societies has become a new paradigm in the emerging field of computational social science. A natural application of this would be the prediction of the society's reaction to a new product in the sense of popularity and adoption rate. However, bridging the gap between "real time monitoring" and "early predicting" remains a big challenge. Here we report on an endeavor to build a minimalistic predictive model for the financial success of movies based on collective activity data of online users. We show that the popularity of a movie can be predicted much before its release by measuring and analyzing the activity level of editors and viewers of the corresponding entry to the movie in Wikipedia, the well-known online encyclopedia.
13 pages, Including Supporting Information, 7 Figures, Download the dataset from: http://wwm.phy.bme.hu/SupplementaryDataS1.zip
References in corpus (5)
- "I Wanted to Predict Elections with Twitter and all I got was this Lousy Paper" -- A Balanced Survey on Election Prediction using Twitter Data
- Opinions, Conflicts and Consensus: Modeling Social Dynamics in a Collaborative Environment
- The statistical laws of popularity: Universal properties of the box office dynamics of motion pictures
- Quantitative analysis of the evolution of novelty in cinema through crowdsourced keywords
- Value production in a collaborative environment
Cited by in corpus (31)
- Quantifying the Effect of Sentiment on Information Diffusion in Social Media
- Global disease monitoring and forecasting with Wikipedia
- Early Predictions of Movie Success: the Who, What, and When of Profitability
- The distorted mirror of Wikipedia: a quantitative analysis of Wikipedia coverage of academics
- The production of information in the attention economy
- Manipulation and abuse on social media
- Using four different online media sources to forecast the crude oil price
- Quantifying and predicting success in show business
- Twitter-based analysis of the dynamics of collective attention to political parties
- Quantitative analysis of the evolution of novelty in cinema through crowdsourced keywords
- Value production in a collaborative environment
- Style in the Age of Instagram: Predicting Success within the Fashion Industry using Social Media
- Wikipedia traffic data and electoral prediction: towards theoretically informed models
- Can electoral popularity be predicted using socially generated big data?
- Understanding Editing Behaviors in Multilingual Wikipedia
- Identifying long-term periodic cycles and memories of collective emotion in online social media
- Self-organization on social media: endo-exo bursts and baseline fluctuations
- Evolving Collaboration, Dependencies, and Use in the Rust Open Source Software Ecosystem
- Dynamics and Biases of Online Attention: The Case of Aircraft Crashes
- Inspiration, Captivation, and Misdirection: Emergent Properties in Networks of Online Navigation
- TwitterPaul: Extracting and Aggregating Twitter Predictions
- A Graph-structured Dataset for Wikipedia Research
- The quantitative measure and statistical distribution of fame
- Memory Remains: Understanding Collective Memory in the Digital Age
- Terrorist attacks sharpen the binary perception of "Us" vs. "Them"
- On the Dynamics of Social Media Popularity: A YouTube Case Study
- Detecting and Gauging Impact on Wikipedia Page Views
- Hybrid Machine Learning Approach to Popularity Prediction of Newly Released Contents for Online Video Streaming Service
- Upscaling human activity data: an ecological perspective
- Social media self-branding and success: Quantitative evidence from a model competition
- Collective Attention towards Scientists and Research Topics