Co-factor analysis of citation networks
arXiv:2408.14604 · doi:10.1080/10618600.2024.2394464
Abstract
One compelling use of citation networks is to characterize papers by their relationships to the surrounding literature. We propose a method to characterize papers by embedding them into two distinct "co-factor" spaces: one describing how papers send citations, and the other describing how papers receive citations. This approach presents several challenges. First, older documents cannot cite newer documents, and thus it is not clear that co-factors are even identifiable. We resolve this challenge by developing a co-factor model for asymmetric adjacency matrices with missing lower triangles and showing that identification is possible. We then frame estimation as a matrix completion problem and develop a specialized implementation of matrix completion because prior implementations are memory bound in our setting. Simulations show that our estimator has promising finite sample properties, and that naive approaches fail to recover latent co-factor structure. We leverage our estimator to investigate 255,780 papers published in statistics journals from 1898 to 2024, resulting in the most comprehensive topic model of the statistics literature to date. We find interpretable co-factors corresponding to many statistical subfields, including time series, variable selection, spatial methods, graphical models, GLM(M)s, causal inference, multiple testing, quantile regression, semiparametrics, dimension reduction, and several more.
Update pre-print to post-print
References in corpus (20)
- Stochastic blockmodels and community structure in networks
- Mixed membership stochastic blockmodels
- Consistency of spectral clustering in stochastic block models
- Fast community detection by SCORE
- Matrix Completion and Low-Rank SVD via Fast Alternating Least Squares
- A network approach to topic models
- Efficiently inferring community structure in bipartite networks
- Coauthorship and Citation Networks for Statisticians
- Noisy low-rank matrix completion with general sampling distribution
- Statistical inference on random dot product graphs: a survey
- Minimax sparse principal subspace estimation in high dimensions
- Clustering Partially Observed Graphs via Convex Optimization
- Reconstructing networks with unknown and heterogeneous errors
- On a 'Two Truths' Phenomenon in Spectral Graph Clustering
- Preferential attachment of communities: the same principle, but a higher level
- Revisiting the Nystrom Method for Improved Large-Scale Machine Learning
- Co-clustering separately exchangeable network data
- Community Detection in Bipartite Networks with Stochastic Blockmodels
- Node Embeddings and Exact Low-Rank Representations of Complex Networks
- Directed mixed membership stochastic blockmodel