The precision of the arithmetic mean, geometric mean and percentiles for citation data: An experimental simulation modelling approach
arXiv:1512.01688 · doi:10.1016/j.joi.2015.12.001
Abstract
When comparing the citation impact of nations, departments or other groups of researchers within individual fields, three approaches have been proposed: arithmetic means, geometric means, and percentage in the top X%. This article compares the precision of these statistics using 97 trillion experimentally simulated citation counts from 6875 sets of different parameters (although all having the same scale parameter) based upon the discretised lognormal distribution with limits from 1000 repetitions for each parameter set. The results show that the geometric mean is the most precise, closely followed by the percentage of a country's articles in the top 50% most cited articles for a field, year and document type. Thus the geometric mean citation count is recommended for future citation-based comparisons between nations. The percentage of a country's articles in the top 1% most cited is a particularly imprecise indicator and is not recommended for international comparisons based on individual fields. Moreover, whereas standard confidence interval formulae for the geometric mean appear to be accurate, confidence interval formulae are less accurate and consistent for percentile indicators. These recommendations assume that the scale parameters of the samples are the same but the choice of indicator is complex and partly conceptual if they are not.
Thelwall, M. (in press). The precision of the arithmetic mean, geometric mean and percentiles for citation data: An experimental simulation modelling approach. Journal of Informetrics
References in corpus (10)
- Power-law distributions in empirical data
- Universality of citation distributions: towards an objective measure of scientific impact
- Tweets vs. Mendeley readers: How do these two social media metrics differ?
- Regression for citation data: An evaluation of different methods
- On the calculation of percentile-based bibliometric indicators
- The relationship between the number of authors of a publication, its citations and the impact factor of the publishing journal: Evidence from Italy
- The VQR, Italy's second national research assessment: Methodological failures and ranking distortions
- Distributions for cited articles from individual subjects and years
- National research impact indicators from Mendeley readers
- National, disciplinary and temporal variations in the extent to which articles with more authors have more impact: Evidence from a geometric field normalised citation indicator
Cited by in corpus (6)
- Three practical field normalised alternative indicator formulae for research evaluation
- The discretised lognormal and hooked power law distributions for complete citation data: Best options for modelling and regression
- Are there too many uncited articles? Zero inflated variants of the discretised lognormal and hooked power law distributions
- The inconsistency of h-index: a mathematical analysis
- Conceptual difficulties in the use of statistical inference in citation analysis
- Evidence for studying interactions between science and policy: An exploration of scholarly and policy references in Overton-indexed policy documents