computer vision

Data Safety: Synthetic Data Quality Analysis Using CIFAKE Dataset

arXiv:2607.12165

summary

The paper examines how synthetic images generated by different methods differ from real images in feature space, color statistics, and model training, and proposes strategies for evaluating and safely mixing synthetic data with real data for image classification.

Abstract

Recently, the societal implementation of high-performance image classification models has expanded rapidly. While these models require vast amounts of training data to improve performance, securing sufficient real images is often impractical. As a means to compensate for this shortage, the use of synthetic data is becoming widespread. However, synthetic images are not necessarily equivalent to real images for training purposes. This study systematically analyzes the differences between two types of synthetic images created by different generation methods and real images from three perspectives: high-dimensional feature space, low-level statistics in color space, and the model training process. Furthermore, it experimentally verifies how synthetic data should be utilized by considering realistic data mixing scenarios. This enables the proposal of an evaluation and application strategy for performing preliminary assessments on synthetic images of unknown quality and safely incorporating them into training. This research aims to contribute to enhancing the reliability and safety of image classification models utilizing synthetic images.

Topics & keywords

#synthetic data#image classification#data quality assessment#feature space analysis#data mixingsynthetic image generationhigh-dimensional feature analysiscolor statisticsCIFAKE datasettraining data augmentation