1 paper · 1 filter
Aaryan Sharma, Vishak Prasad C, Virendra Singh +1
Vision-Language Models (VLMs) are highly effective in retrieving semantically relevant images. However, in practice, relevance alone is often insufficient. Systems must also achiev…