1 paper · 1 filter
Gokul Karthik Kumar, Iheb Chaabane, Kebin Wu
Vision-language models (VLMs) excel in various visual benchmarks but are often constrained by the lack of high-quality visual fine-tuning data. To address this challenge, we introd…