Veridical Data Science
arXiv:1901.08152 · doi:10.1073/pnas.1901326117
Abstract
Building and expanding on principles of statistics, machine learning, and scientific inquiry, we propose the predictability, computability, and stability (PCS) framework for veridical data science. Our framework, comprised of both a workflow and documentation, aims to provide responsible, reliable, reproducible, and transparent results across the entire data science life cycle. The PCS workflow uses predictability as a reality check and considers the importance of computation in data collection/storage and algorithm design. It augments predictability and computability with an overarching stability principle for the data science life cycle. Stability expands on statistical uncertainty considerations to assess how human judgment calls impact data results through data and model/algorithm perturbations. Moreover, we develop inference procedures that build on PCS, namely PCS perturbation intervals and PCS hypothesis testing, to investigate the stability of data results relative to problem formulation, data cleaning, modeling decisions, and interpretations. We illustrate PCS inference through neuroscience and genomics projects of our own and others and compare it to existing methods in high dimensional, sparse linear model simulations. Over a wide range of misspecified simulation models, PCS inference demonstrates favorable performance in terms of ROC curves. Finally, we propose PCS documentation based on R Markdown or Jupyter Notebook, with publicly available, reproducible codes and narratives to back up human choices made throughout an analysis. The PCS workflow and documentation are demonstrated in a genomics case study available on Zenodo.
References in corpus (3)
Cited by in corpus (25)
- Evaluating probabilistic classifiers: Reliability diagrams and score decompositions revisited
- Anchor regression: heterogeneous data meets causality
- Principles for data analysis workflows
- Regression Diagnostics meets Forecast Evaluation: Conditional Calibration, Reliability Diagrams, and Coefficient of Determination
- Interpretable Random Forests via Rule Extraction
- Learning stable and predictive structures in kinetic systems: Benefits of a causal approach
- Landscape of R packages for eXplainable Artificial Intelligence
- Designing Reinforcement Learning Algorithms for Digital Interventions: Pre-implementation Guidelines
- Federated Accelerated Stochastic Gradient Descent
- Honest calibration assessment for binary outcome predictions
- Provable Boolean Interaction Recovery from Tree Ensemble obtained via Random Forests
- Potential for allocative harm in an environmental justice data tool
- Controlling False Discovery Rate Using Gaussian Mirrors
- Impact of Accuracy on Model Interpretations
- A picture guide to cancer progression and monotonic accumulation models: evolutionary assumptions, plausible interpretations, and alternative uses
- Next Waves in Veridical Network Embedding
- Stable discovery of interpretable subgroups via calibration in causal studies
- Learning from learning machines: a new generation of AI technology to meet the needs of science
- Deconfounding and Causal Regularization for Stability and External Validity
- The Shapley Value of coalition of variables provides better explanations
- Revisiting minimum description length complexity in overparameterized models
- Neural Gaussian Mirror for Controlled Feature Selection in Neural Networks
- Rejoinder: Learning Optimal Distributionally Robust Individualized Treatment Rules
- Contrasting pre-vaccine COVID-19 waves in Italy through Functional Data Analysis
- Comments on Leo Breiman's paper 'Statistical Modeling: The Two Cultures' (Statistical Science, 2001, 16(3), 199-231)