Caveats for using statistical significance tests in research assessments
arXiv:1112.2516 · doi:10.1016/j.joi.2012.08.005
Abstract
This paper raises concerns about the advantages of using statistical significance tests in research assessments as has recently been suggested in the debate about proper normalization procedures for citation indicators. Statistical significance tests are highly controversial and numerous criticisms have been leveled against their use. Based on examples from articles by proponents of the use statistical significance tests in research assessments, we address some of the numerous problems with such tests. The issues specifically discussed are the ritual practice of such tests, their dichotomous application in decision making, the difference between statistical and substantive significance, the implausibility of most null hypotheses, the crucial assumption of randomness, as well as the utility of standard errors and confidence intervals for inferential purposes. We argue that applying statistical significance tests and mechanically adhering to their results is highly problematic and detrimental to critical thinking. We claim that the use of such tests do not provide any advantages in relation to citation indicators, interpretations of them, or the decision making processes based upon them. On the contrary their use may be harmful. Like many other critics, we generally believe that statistical significance tests are over- and misused in the social sciences including scientometrics and we encourage a reform on these matters.
Accepted version for Journal of Informetrics
References in corpus (1)
Cited by in corpus (7)
- A review of the characteristics of 108 author-level bibliometric indicators
- The precision of the arithmetic mean, geometric mean and percentiles for citation data: An experimental simulation modelling approach
- A critical cluster analysis of 44 indicators of author-level performance
- Conceptual difficulties in the use of statistical inference in citation analysis
- Data inaccuracy quantification and uncertainty propagation for bibliometric indicators
- Funnel plots for visualizing uncertainty in the research performance of institutions
- Confidence intervals for normalised citation counts: Can they delimit underlying research capability?