2 citations · 4 across the 3 of their papers we have counts for
4 papers
Constructing Phrase-level Semantic Labels to Form Multi-Grained Supervision for Image-Text Retrieval
Zhihao Fan, Zhongyu Wei, Zejun Li +4
Existing research for image text retrieval mainly relies on sentence-level supervision to distinguish matched and mismatched sentences for a query image. However, semantic mismatch…
TCIC: Theme Concepts Learning Cross Language and Vision for Image Captioning
Zhihao Fan, Zhongyu Wei, Siyuan Wang +4
Existing research for image captioning usually represents an image using a scene graph with low-level facts (objects and relations) and fails to capture the high-level semantics. I…
An Unsupervised Sampling Approach for Image-Sentence Matching Using Document-Level Structural Information
Zejun Li, Zhongyu Wei, Zhihao Fan +2
In this paper, we focus on the problem of unsupervised image-sentence matching. Existing research explores to utilize document-level structural information to sample positive and n…
Keep it Consistent: Topic-Aware Storytelling from an Image Stream via Iterative Multi-agent Communication
Ruize Wang, Zhongyu Wei, Ying Cheng +5
Visual storytelling aims to generate a narrative paragraph from a sequence of images automatically. Existing approaches construct text description independently for each image and…