Data Science vs. Statistics: Two Cultures?
arXiv:1801.00371 · doi:10.1007/s42081-018-0009-3
Abstract
Data science is the business of learning from data, which is traditionally the business of statistics. Data science, however, is often understood as a broader, task-driven and computationally-oriented version of statistics. Both the term data science and the broader idea it conveys have origins in statistics and are a reaction to a narrower view of data analysis. Expanding upon the views of a number of statisticians, this paper encourages a big-tent view of data analysis. We examine how evolving approaches to modern data analysis relate to the existing discipline of statistics (e.g. exploratory analysis, machine learning, reproducibility, computation, communication and the role of theory). Finally, we discuss what these trends mean for the future of statistics by highlighting promising directions for communication, education and research.
References in corpus (8)
- Towards A Rigorous Science of Interpretable Machine Learning
- Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings
- Classifier Technology and the Illusion of Progress
- "Why Should I Trust You?": Explaining the Predictions of Any Classifier
- Curriculum Guidelines for Undergraduate Programs in Data Science
- Reproducible Research Can Still Be Wrong: Adopting a Prevention Approach
- Object oriented data analysis: Sets of trees
- Differentially Private Gaussian Processes
Cited by in corpus (4)
- Teaching Responsible Data Science: Charting New Pedagogical Territory
- The algebra and machine representation of statistical models
- The future of urban models in the Big Data and AI era: a bibliometric analysis (2000-2019)
- Toward a Knowledge Discovery Framework for Data Science Job Market in the United States