1 paper
Jiacong Wang, Bohong Wu, Haiyong Jiang +4
Recent advances in Vision-Language Models (VLMs) and the scarcity of high-quality multi-modal alignment data have inspired numerous researches on synthetic VLM data generation. The…