1 paper · 1 filter
Siyi Liu, Xiaorong Zhu, Enjun Du +6
Vision-language retrieval with CLIP-style dual encoders achieves strong cross-modal performance, yet practical accuracy often hinges on localized semantic distinctions where top-ra…