Consistencies and inconsistencies between model selection and link prediction in networks
arXiv:1705.07967 · doi:10.1103/PhysRevE.97.062316
Abstract
A principled approach to understand network structures is to formulate generative models. Given a collection of models, however, an outstanding key task is to determine which one provides a more accurate description of the network at hand, discounting statistical fluctuations. This problem can be approached using two principled criteria that at first may seem equivalent: selecting the most plausible model in terms of its posterior probability; or selecting the model with the highest predictive performance in terms of identifying missing links. Here we show that while these two approaches yield consistent results in most of cases, there are also notable instances where they do not, that is, where the most plausible model is not the most predictive. We show that in the latter case the improvement of predictive performance can in fact lead to overfitting both in artificial and empirical settings. Furthermore, we show that, in general, the predictive performance is higher when we average over collections of models that are individually less plausible, than when we consider only the single most plausible model.
12 pages, 6 figures, 1 table
References in corpus (9)
- Modularity and community structure in networks
- Hierarchical structure and the prediction of missing links in networks
- Stochastic blockmodels and community structure in networks
- An information-theoretic framework for resolving community structure in complex networks
- Missing and spurious interactions and the reconstruction of complex networks
- Parsimonious module inference in large networks
- Community detection, link prediction, and layer interdependence in multilayer networks
- Predicting human preferences using the block structure of complex social networks
- Algorithmic detectability threshold of the stochastic block model
Cited by in corpus (23)
- A network approach to topic models
- A Review of Stochastic Block Models and Extensions for Graph Clustering
- Evaluating Overfit and Underfit in Models of Network Community Structure
- Stacking Models for Nearly Optimal Link Prediction in Complex Networks
- Progresses and Challenges in Link Prediction
- Bayesian stochastic blockmodeling
- Descriptive vs. inferential community detection in networks: pitfalls, myths, and half-truths
- Link prediction with hyperbolic geometry
- Review on Learning and Extracting Graph Features for Link Prediction
- Community Detection in Bipartite Networks with Stochastic Blockmodels
- Disentangling homophily, community structure and triadic closure in networks
- Human mobility is well described by closed-form gravity-like models learned automatically from data
- Fundamental limits to learning closed-form mathematical models from data
- Tensorial and bipartite block models for link prediction in layered networks and temporal networks
- Multilayer Networks for Text Analysis with Multiple Data Types
- Mapping flows on sparse networks with missing links
- Local-ring network automata and the impact of hyperbolic geometry in complex network link-prediction
- Optimal prediction of decisions and model selection in social dilemmas using block models
- Implicit models, latent compression, intrinsic biases, and cheap lunches in community detection
- Information Evolution in Complex Networks
- Link Prediction Accuracy on Real-World Networks Under Non-Uniform Missing Edge Patterns
- Description length of canonical and microcanonical models
- Synthetic graphs for link prediction benchmarking