1 paper · 1 filter
Diji Yang, Minghao Liu, Chung-Hsiang Lo +2
Vision-language models (VLMs) have shown strong performance on text-to-image retrieval benchmarks. However, bridging this success to real-world applications remains a challenge. In…