1 paper
Rui Cai, Zhiyu Dong, Jianfeng Dong +1
Existing cross-modal retrieval methods typically rely on large-scale vision-language pair data. This makes it challenging to efficiently develop a cross-modal retrieval model for u…