1 paper
Aaryan Sharma, Vishak Prasad C, Virendra Singh +1
Vision-Language Models (VLMs) are highly effective in retrieving semantically relevant images. However, in practice, relevance alone is often insufficient. Systems must also achiev…