Sampling in Software Engineering Research: A Critical Review and Guidelines
arXiv:2002.07764
Abstract
Representative sampling appears rare in empirical software engineering research. Not all studies need representative samples, but a general lack of representative sampling undermines a scientific field. This article therefore reports a critical review of the state of sampling in recent, high-quality software engineering research. The key findings are: (1) random sampling is rare; (2) sophisticated sampling strategies are very rare; (3) sampling, representativeness and randomness often appear misunderstood. These findings suggest that software engineering research has a generalizability crisis. To address these problems, this paper synthesizes existing knowledge of sampling into a succinct primer and proposes extensive guidelines for improving the conduct, presentation and evaluation of sampling in software engineering research. It is further recommended that while researchers should strive for more representative samples, disparaging non-probability sampling is generally capricious and particularly misguided for predominately qualitative research.
38 pages, 8 tables, accepted for publication in Empirical Software Engineering
Cited by in corpus (7)
- Predictors of Well-being and Productivity among Software Professionals during the COVID-19 Pandemic -- A Longitudinal Study
- A Large-Scale Security-Oriented Static Analysis of Python Packages in PyPI
- The Mind Is a Powerful Place: How Showing Code Comprehensibility Metrics Influences Code Understanding
- Use and Misuse of the Term Experiment in Mining Software Repositories Research
- Not All Requirements Prioritization Criteria Are Equal at All Times: A Quantitative Analysis
- Towards a Methodology for Participant Selection in Software Engineering Experiments. A Vision of the Future
- No Silver Bullets: Why Understanding Software Cycle Time is Messy, Not Magic