Regression for citation data: An evaluation of different methods
arXiv:1510.08877 · doi:10.1016/j.joi.2014.09.011
Abstract
Citations are increasingly used for research evaluations. It is therefore important to identify factors affecting citation scores that are unrelated to scholarly quality or usefulness so that these can be taken into account. Regression is the most powerful statistical technique to identify these factors and hence it is important to identify the best regression strategy for citation data. Citation counts tend to follow a discrete lognormal distribution and, in the absence of alternatives, have been investigated with negative binomial regression. Using simulated discrete lognormal data (continuous lognormal data rounded to the nearest integer) this article shows that a better strategy is to add one to the citations, take their log and then use the general linear (ordinary least squares) model for regression (e.g., multiple linear regression, ANOVA), or to use the generalised linear model without the log. Reasonable results can also be obtained if all the zero citations are discarded, the log is taken of the remaining citation counts and then the general linear model is used, or if the generalised linear model is used with the continuous lognormal distribution. Similar approaches are recommended for altmetric data, if it proves to be lognormally distributed.
References in corpus (5)
- Universality of citation distributions: towards an objective measure of scientific impact
- How well developed are altmetrics? A cross-disciplinary analysis of the presence of 'alternative metrics' in scientific publications
- National research assessment exercises: a comparison of peer review and bibliometrics rankings
- Distributions for cited articles from individual subjects and years
- National peer-review research assessment exercises for the hard sciences can be a complete waste of money: the Italian case
Cited by in corpus (20)
- The citation advantage of linking publications to research data
- Do altmetrics correlate with the quality of papers? A large-scale empirical study based on F1000Prime data
- Could scientists use Altmetric.com scores to predict longer term citation counts?
- The discretised lognormal and hooked power law distributions for complete citation data: Best options for modelling and regression
- Preprints as accelerator of scholarly communication: An empirical analysis in Mathematics
- On the relationships between bibliographic characteristics of scientific documents and citation and Mendeley readership counts: A large-scale analysis of Web of Science publications
- The precision of the arithmetic mean, geometric mean and percentiles for citation data: An experimental simulation modelling approach
- Science and Facebook: the same popularity law!
- National, disciplinary and temporal variations in the extent to which articles with more authors have more impact: Evidence from a geometric field normalised citation indicator
- The citation advantage of foreign language references for Chinese social science papers
- The funding effect on citation and social attention: the UN Sustainable Development Goals (SDGs) as a case study
- Citation count distributions for large monodisciplinary journals
- Do more heads imply better performance? An empirical study of team thought leaders' impact on scientific team performance
- The inconsistency of h-index: a mathematical analysis
- Research assessment by percentile-based double rank analysis
- The Accuracy of Confidence Intervals for Field Normalised Indicators
- Biblioranking fundamental physics
- The role of preprints in open science: Accelerating knowledge transfer from science to technology
- On The Peer Review Reports: Does Size Matter?
- Degree distributions in networks: beyond the power law