Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
Let's Go Shopping (LGS) -- Web-Scale Image-Text Dataset for Visual Concept Understanding
Yatong Bai, Utsav Garg, Apaar Shanker +10
Vision and vision-language applications of neural networks, such as image classification and captioning, rely on large-scale annotated datasets that require non-trivial data-collec…
cs.CV2022
Discriminative Supervised Subspace Learning for Cross-modal Retrieval
Haoming Zhang, Xiao-Jun Wu, Tianyang Xu +1
Nowadays the measure between heterogeneous data is still an open problem for cross-modal retrieval. The core of cross-modal retrieval is how to measure the similarity between diffe…