9 citations · 13 across the 4 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2022★ 1 cited
Dissecting Deep Metric Learning Losses for Image-Text Retrieval
Hong Xuan, Xi Chen
Visual-Semantic Embedding (VSE) is a prevalent approach in image-text retrieval by learning a joint embedding space between the image and language modalities where semantic similar…
cs.CV2022★ 3 cited
All You May Need for VQA are Image Captions
Soravit Changpinyo, Doron Kukliansky, Idan Szpektor +3
Visual Question Answering (VQA) has benefited from increasingly sophisticated models, but has not enjoyed the same level of engagement in terms of data creation. In this paper, we…