29 citations · 65 across the 5 of their papers we have counts for
5 papers
Learning Text-Image Joint Embedding for Efficient Cross-Modal Retrieval with Deep Feature Engineering
Zhongwei Xie, Ling Liu, Yanzhao Wu +2
This paper introduces a two-phase deep feature engineering framework for efficient learning of semantics enhanced joint embedding, which clearly separates the deep feature engineer…
Visual-aware Attention Dual-stream Decoder for Video Captioning
Zhixin Sun, Xian Zhong, Shuqin Chen +2
Video captioning is a challenging task that captures different visual parts and describes them in sentences, for it requires visual and linguistic coherence. The attention mechanis…
Learning Joint Embedding with Modality Alignments for Cross-Modal Retrieval of Recipes and Food Images
Zhongwei Xie, Ling Liu, Lin Li +1
This paper presents a three-tier modality alignment approach to learning text-image joint embedding, coined as JEMA, for cross-modal retrieval of cooking recipes and food images. T…
Efficient Deep Feature Calibration for Cross-Modal Joint Embedding Learning
Zhongwei Xie, Ling Liu, Lin Li +1
This paper introduces a two-phase deep feature calibration framework for efficient learning of semantics enhanced text-image cross-modal joint embedding, which clearly separates th…
Learning TFIDF Enhanced Joint Embedding for Recipe-Image Cross-Modal Retrieval Service
Zhongwei Xie, Ling Liu, Yanzhao Wu +2
It is widely acknowledged that learning joint embeddings of recipes with images is challenging due to the diverse composition and deformation of ingredients in cooking procedures.…