A Scoping Review of Synthetic Data Generation by Language Models in Biomedical Research and Application: Data Utility and Quality Perspectives
arXiv:2506.16594 · doi:10.1007/s41666-026-00229-9
Abstract
Synthetic data generation using large language models (LLMs) demonstrates substantial promise in addressing biomedical data challenges and shows increasing adoption in biomedical research. This study systematically reviews recent advances in synthetic data generation for biomedical applications and clinical research, focusing on how LLMs address data scarcity, utility, and quality issues with different modalities. We conducted a scoping review following PRISMA-ScR guidelines and searched literature published between 2020 and 2025 through PubMed, ACM, Web of Science, and Google Scholar. A total of 59 studies were included based on relevance to synthetic data generation in biomedical contexts. Among the reviewed studies, the predominant data modalities were unstructured texts (78.0\%), tabular data (13.6\%), and multimodal sources (8.4\%). Common generation methods included LLM prompting (74.6\%), fine-tuning (20.3\%), and specialized models (5.1\%). Evaluations were heterogeneous: intrinsic metrics (27.1\%), human-in-the-loop assessments (44.1\%), and LLM-based evaluations (13.6\%). However, limitations and key barriers persist in data modalities, domain utility, resource and model accessibility, and standardized evaluation protocols. Future efforts may focus on developing standardized, transparent evaluation frameworks and expanding accessibility to support effective applications in biomedical research.
References in corpus (9)
- The Role of AI in Drug Discovery: Challenges, Opportunities, and Strategies
- A Study of Generative Large Language Model for Medical Research and Healthcare
- Socially Aware Synthetic Data Generation for Suicidal Ideation Detection Using Large Language Models
- Synthesize High-dimensional Longitudinal Electronic Health Records via Hierarchical Autoregressive Language Model
- De-identification is not enough: a comparison between de-identified and synthetic clinical notes
- Two Directions for Clinical Data Generation with Large Language Models: Data-to-Label and Label-to-Data
- Textual Data Augmentation for Patient Outcomes Prediction
- NoteChat: A Dataset of Synthetic Doctor-Patient Conversations Conditioned on Clinical Notes
- Robust Privacy Amidst Innovation with Large Language Models Through a Critical Assessment of the Risks