most citedLearning TFIDF Enhanced Joint Embedding for Recipe-Image Cross-Modal Retrieval Service

29 citations · 65 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV202124 cited

Learning Text-Image Joint Embedding for Efficient Cross-Modal Retrieval with Deep Feature Engineering

Zhongwei Xie, Ling Liu, Yanzhao Wu +2

This paper introduces a two-phase deep feature engineering framework for efficient learning of semantics enhanced joint embedding, which clearly separates the deep feature engineer…

cs.CV2021

Visual-aware Attention Dual-stream Decoder for Video Captioning

Zhixin Sun, Xian Zhong, Shuqin Chen +2

Video captioning is a challenging task that captures different visual parts and describes them in sentences, for it requires visual and linguistic coherence. The attention mechanis…

cs.CV20219 cited

Learning Joint Embedding with Modality Alignments for Cross-Modal Retrieval of Recipes and Food Images

Zhongwei Xie, Ling Liu, Lin Li +1

This paper presents a three-tier modality alignment approach to learning text-image joint embedding, coined as JEMA, for cross-modal retrieval of cooking recipes and food images. T…

cs.CV20213 cited

Efficient Deep Feature Calibration for Cross-Modal Joint Embedding Learning

Zhongwei Xie, Ling Liu, Lin Li +1

This paper introduces a two-phase deep feature calibration framework for efficient learning of semantics enhanced text-image cross-modal joint embedding, which clearly separates th…

cs.CV202129 cited

Learning TFIDF Enhanced Joint Embedding for Recipe-Image Cross-Modal Retrieval Service

Zhongwei Xie, Ling Liu, Yanzhao Wu +2

It is widely acknowledged that learning joint embeddings of recipes with images is challenging due to the diverse composition and deformation of ingredients in cooking procedures.…