Methods for generating and evaluating synthetic longitudinal patient data: a systematic review
arXiv:2309.12380 · doi:10.1007/s41666-025-00223-7
Abstract
The rapid growth in data availability has facilitated research and development, yet not all industries have benefited equally due to legal and privacy constraints. The healthcare sector faces significant challenges in utilizing patient data because of concerns about data security and confidentiality. To address this, various privacy-preserving methods, including synthetic data generation, have been proposed. Synthetic data replicate existing data as closely as possible, acting as a proxy for sensitive information. While patient data are often longitudinal, this aspect remains underrepresented in existing reviews of synthetic data generation in healthcare. This paper maps and describes methods for generating and evaluating synthetic longitudinal patient data in real-life settings through a systematic literature review, conducted following the PRISMA guidelines and incorporating data from five databases up to May 2024. Thirty-nine methods were identified, with four addressing all challenges of longitudinal data generation, though none included privacy-preserving mechanisms. Resemblance was evaluated in most studies, utility in the majority, and privacy in just over half. Only a small fraction of studies assessed all three aspects. Our findings highlight the need for further research in this area.
References in corpus (11)
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Generative AI for Synthetic Data Across Multiple Medical Modalities: A Systematic Review of Recent Developments and Challenges
- A Multifaceted Benchmarking of Synthetic Electronic Health Record Generation Models
- Machine Learning for Synthetic Data Generation: A Review
- A review of Generative Adversarial Networks for Electronic Health Records: applications, evaluation measures and data sources
- Synthesize High-dimensional Longitudinal Electronic Health Records via Hierarchical Autoregressive Language Model
- Synthcity: facilitating innovative use cases of synthetic data in different data modalities
- Sharing is CAIRing: Characterizing Principles and Assessing Properties of Universal Privacy Evaluation for Synthetic Tabular Data
- CEHR-GPT: Generating Electronic Health Records with Chronological Patient Timelines
- Continual Release of Differentially Private Synthetic Data from Longitudinal Data Collections
- Collaborative Synthesis of Patient Records through Multi-Visit Health State Inference