1 paper
Chang Kong, Yuebing Li, Peng Mo +2
The scarcity of hard negative samples in current vision-language datasets significantly hinders fine-grained perception. To address this, we propose FineGen, a VLM-based Multi-Agen…