Synthetic Data Generation Using Large Language Models: Advances in Text and Code
arXiv:2503.14023 · doi:10.1109/ACCESS.2025.3589503
Abstract
This survey reviews how large language models (LLMs) are transforming synthetic training data generation in both natural language and code domains. By producing artificial but task-relevant examples, these models can significantly augment or even substitute for real-world datasets, particularly in scenarios where labeled data is scarce, expensive, or sensitive. This paper surveys recent advances in leveraging LLMs to create synthetic text and code, highlighting key techniques such as prompt-based generation, retrieval-augmented pipelines, and iterative self-refinement. We examine how these methods can enrich low-resource tasks (e.g., classification, question answering) and facilitate code-centric applications (e.g., instruction tuning, code translation, bug repair) through automated verification of functional correctness. Alongside potential benefits - cost-effectiveness, broad coverage, and controllable diversity - we discuss the accompanying challenges, including factual inaccuracies in generated text, insufficient stylistic or distributional realism, and risks of bias amplification. Proposed mitigation strategies range from filtering and weighting synthetic outputs to reinforcement learning with execution feedback in code domains. We conclude by outlining open research directions, such as automated prompt engineering, cross-modal data synthesis, and robust evaluation frameworks, underscoring the growing importance of LLM-generated synthetic data in accelerating AI development while emphasizing ethical and quality safeguards.
24 pages, 6 tables, 1 figure, 64 references
References in corpus (23)
- Distilling the Knowledge in a Neural Network
- Evaluating Large Language Models Trained on Code
- Reflexion: Language Agents with Verbal Reinforcement Learning
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
- Measuring Coding Challenge Competence With APPS
- InCoder: A Generative Model for Code Infilling and Synthesis
- DiffEdit: Diffusion-based semantic image editing with mask guidance
- WizardCoder: Empowering Code Large Language Models with Evol-Instruct
- Effective Data Augmentation With Diffusion Models
- A Survey of Text-to-SQL in the Era of LLMs: Where are we, and where are we going?
- PanGu-Coder: Program Synthesis with Function-Level Language Modeling
- Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning
- Language Models Can Teach Themselves to Program Better
- Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
- AixBench: A Code Generation Benchmark Dataset
- InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback
- A Survey on Data Synthesis and Augmentation for Large Language Models
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
- Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities
- LLM-Assisted Code Cleaning For Training Accurate Code Generators
- Data Generation Using Large Language Models for Text Classification: An Empirical Case Study
- GSURE-Based Diffusion Model Training with Corrupted Data
- Training Language Models on Synthetic Edit Sequences Improves Code Synthesis