23 citations · 59 across the 6 of their papers we have counts for
6 papers
Position Focused Attention Network for Image-Text Matching
Yaxiong Wang, Hao Yang, Xueming Qian +4
Image-text matching tasks have recently attracted a lot of attention in the computer vision field. The key point of this cross-domain problem is how to accurately measure the simil…
Hallucinating Optical Flow Features for Video Classification
Yongyi Tang, Lin Ma, Lianqiang Zhou
Appearance and motion are two key components to depict and characterize the video content. Currently, the two-stream models have achieved state-of-the-art performances on video cla…
Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video
Zhenfang Chen, Lin Ma, Wenhan Luo +1
In this paper, we address a novel task, namely weakly-supervised spatio-temporally grounding natural sentence in video. Specifically, given a natural sentence and a video, we local…
Reconstruct and Represent Video Contents for Captioning via Reinforcement Learning
Wei Zhang, Bairui Wang, Lin Ma +1
In this paper, the problem of describing visual contents of a video sequence with natural language is addressed. Unlike previous video captioning work mainly exploiting the cues of…
Spatio-temporal Video Re-localization by Warp LSTM
Yang Feng, Lin Ma, Wei Liu +1
The need for efficiently finding the video content a user wants is increasing because of the erupting of user-generated videos on the Web. Existing keyword-based or content-based v…
Image Deformation Meta-Networks for One-Shot Learning
Zitian Chen, Yanwei Fu, Yu-Xiong Wang +3
Humans can robustly learn novel visual concepts even when images undergo various deformations and lose certain information. Mimicking the same behavior and synthesizing deformed in…