10 Simple Rules for the Care and Feeding of Scientific Data
arXiv:1401.2134 · doi:10.1371/journal.pcbi.1003542
Abstract
This article offers a short guide to the steps scientists can take to ensure that their data and associated analyses continue to be of value and to be recognized. In just the past few years, hundreds of scholarly papers and reports have been written on questions of data sharing, data provenance, research reproducibility, licensing, attribution, privacy, and more, but our goal here is not to review that literature. Instead, we present a short guide intended for researchers who want to know why it is important to "care for and feed" data, with some practical advice on how to do that.
Accepted in PLOS Computational Biology. This paper was written collaboratively, on the web, in the open, using Authorea. The living version of this article, which includes sources and history, is available at http://www.authorea.com/3410/
Cited by in corpus (11)
- "Garbage In, Garbage Out" Revisited: What Do Machine Learning Application Papers Report About Human-Labeled Training Data?
- Ontology Development Kit: a toolkit for building, maintaining, and standardising biomedical ontologies
- Dataset Search In Biodiversity Research: Do Metadata In Data Repositories Reflect Scholarly Information Needs?
- Jupyter notebooks as discovery mechanisms for open science: Citation practices in the astronomy community
- LSSGalPy: Interactive Visualization of the Large-scale Environment Around Galaxies
- Interpretable Uncertainty Quantification in AI for HEP
- From Data Processes to Data Products: Knowledge Infrastructures in Astronomy
- Lost Data in Electron Microscopy
- Formalizing Privacy Laws for License Generation and Data Repository Decision Automation
- An Intermediate Data-driven Methodology for Scientific Workflow Management System to Support Reusability
- Garbage In, Garbage Out? Do Machine Learning Application Papers in Social Computing Report Where Human-Labeled Training Data Comes From?