Quantifying and suppressing ranking bias in a large citation network
arXiv:1703.08071 · doi:10.1016/j.joi.2017.05.014
Abstract
It is widely recognized that citation counts for papers from different fields cannot be directly compared because different scientific fields adopt different citation practices. Citation counts are also strongly biased by paper age since older papers had more time to attract citations. Various procedures aim at suppressing these biases and give rise to new normalized indicators, such as the relative citation count. We use a large citation dataset from Microsoft Academic Graph and a new statistical framework based on the Mahalanobis distance to show that the rankings by well known indicators, including the relative citation count and Google's PageRank score, are significantly biased by paper field and age. We propose a general normalization procedure motivated by the -score which produces much less biased rankings when applied to citation count and PageRank score.
Main text (pp. 1-12) and Appendices (pp. 13-17)
References in corpus (5)
- Universality of citation distributions: towards an objective measure of scientific impact
- Finding Scientific Gems with Google
- Ranking Scientific Publications Using a Simple Model of Network Traffic
- Promise and Pitfalls of Extending Google's PageRank Algorithm to Citation Networks
- Identification of milestone papers through time-balanced network centrality
Cited by in corpus (6)
- Ranking in evolving complex networks
- The coverage of Microsoft Academic: Analyzing the publication output of a university
- Early identification of important patents through network centrality
- Unbiased evaluation of ranking metrics reveals consistent performance in science and technology citation data
- The long-term impact of ranking algorithms in growing networks
- Network-based ranking in social systems: three challenges