Recommendations on test datasets for evaluating AI solutions in pathology
arXiv:2204.14226 · doi:10.1038/s41379-022-01147-y
Abstract
Artificial intelligence (AI) solutions that automatically extract information from digital histology images have shown great promise for improving pathological diagnosis. Prior to routine use, it is important to evaluate their predictive performance and obtain regulatory approval. This assessment requires appropriate test datasets. However, compiling such datasets is challenging and specific recommendations are missing. A committee of various stakeholders, including commercial AI developers, pathologists, and researchers, discussed key aspects and conducted extensive literature reviews on test datasets in pathology. Here, we summarize the results and derive general recommendations for the collection of test datasets. We address several questions: Which and how many images are needed? How to deal with low-prevalence subsets? How can potential bias be detected? How should datasets be reported? What are the regulatory requirements in different countries? The recommendations are intended to help AI developers demonstrate the utility of their products and to help regulatory agencies and end users verify reported performance measures. Further research is needed to formulate criteria for sufficiently representative test datasets so that AI solutions can operate with less user intervention and better support diagnostic workflows in the future.
References in corpus (5)
- PanNuke Dataset Extension, Insights and Baselines
- Does Your Dermatology Classifier Know What It Doesn't Know? Detecting the Long-Tail of Unseen Conditions
- A Benchmark of Medical Out of Distribution Detection
- Lizard: A Large-Scale Dataset for Colonic Nuclear Instance Segmentation and Classification
- Confidence-based Out-of-Distribution Detection: A Comparative Study and Analysis
Cited by in corpus (6)
- Applications of artificial intelligence in the analysis of histopathology images of gliomas: a review
- Seeing the random forest through the decision trees. Supporting learning health systems from histopathology with machine learning models: Challenges and opportunities
- Performance of externally validated machine learning models based on histopathology images for the diagnosis, classification, prognosis, or treatment outcome prediction in female breast cancer: A systematic review
- Joining Forces for Pathology Diagnostics with AI Assistance: The EMPAIA Initiative
- The NCI Imaging Data Commons as a platform for reproducible research in computational pathology
- From slides to AI-ready maps: Standardized multi-layer tissue maps as metadata for artificial intelligence in digital pathology