1 paper · 1 filter
Chang Kong, Yuebing Li, Peng Mo +2
The scarcity of hard negative samples in current vision-language datasets significantly hinders fine-grained perception. To address this, we propose FineGen, a VLM-based Multi-Agen…