Benchmarking emergency department triage prediction models with machine learning and large public electronic health records
arXiv:2111.11017 · doi:10.1038/s41597-022-01782-9
Abstract
The demand for emergency department (ED) services is increasing across the globe, particularly during the current COVID-19 pandemic. Clinical triage and risk assessment have become increasingly challenging due to the shortage of medical resources and the strain on hospital infrastructure caused by the pandemic. As a result of the widespread use of electronic health records (EHRs), we now have access to a vast amount of clinical data, which allows us to develop predictive models and decision support systems to address these challenges. To date, however, there are no widely accepted benchmark ED triage prediction models based on large-scale public EHR data. An open-source benchmarking platform would streamline research workflows by eliminating cumbersome data preprocessing, and facilitate comparisons among different studies and methodologies. In this paper, based on the Medical Information Mart for Intensive Care IV Emergency Department (MIMIC-IV-ED) database, we developed a publicly available benchmark suite for ED triage predictive models and created a benchmark dataset that contains over 400,000 ED visits from 2011 to 2019. We introduced three ED-based outcomes (hospitalization, critical outcomes, and 72-hour ED reattendance) and implemented a variety of popular methodologies, ranging from machine learning methods to clinical scoring systems. We evaluated and compared the performance of these methods against benchmark tasks. Our codes are open-source, allowing anyone with MIMIC-IV-ED data access to perform the same steps in data processing, benchmark model building, and experiments. This study provides future researchers with insights, suggestions, and protocols for managing raw data and developing risk triaging tools for emergency care.
References in corpus (4)
- Deep learning for temporal data representation in electronic health records: A systematic review of challenges and methodologies
- A novel interpretable machine learning system to generate clinical risk scores: An application for predicting early mortality or unplanned readmission in a retrospective cohort study
- AutoScore-Survival: Developing interpretable machine learning-based time-to-event scores with right-censored survival data
- AutoScore-Imbalance: An interpretable machine learning tool for development of clinical scores with rare events data
Cited by in corpus (4)
- Federated Learning for Clinical Structured Data: A Benchmark Comparison of Engineering and Statistical Approaches
- Emergency Department Decision Support using Clinical Pseudo-notes
- Developing Federated Time-to-Event Scores Using Heterogeneous Real-World Survival Data
- Towards symbolic regression for interpretable clinical decision scores