paper

From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection

arXiv:2609.38010 · doi:10.1145/3787256.3787261

Abstract

Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data generation paradigms, (1) Unity Simulation-based rendering and (2) Controllable Diffusion-based generation (CIA), for object detection under real data-scarce conditions. A unified experimental framework enables controlled dataset mixing across real, simulated, and generative sources, while maintaining identical model and training settings. Quantitative evaluation using Precision, Recall, mAP, and custom -metrics, reveals that neither simulation nor generative augmentation alone achieves optimal transferability. Unity-only training yields an [email protected] drop of relative to real data, while CIA-only training shows a milder degradation. Hybrid compositions significantly improve performance, with the 90\% real + 10\% Unity configuration achieving the best overall [email protected] of ( over baseline), and the 90\% real + 10\% CIA configuration maximizing precision at . Results demonstrate that limited synthetic inclusion enhances generalization, while excessive substitution induces domain drift.

References in corpus (3)