14 citations · 14 across the 1 of their papers we have counts for
5 papers
Real-Time Referring Expression Comprehension by Single-Stage Grounding Network
Xinpeng Chen, Lin Ma, Jingyuan Chen +3
In this paper, we propose a novel end-to-end model, namely Single-Stage Grounding network (SSG), to localize the referent given a referring expression within an image. Different fr…
Fine-grained Video Attractiveness Prediction Using Multimodal Deep Learning on a Large Real-world Dataset
Xinpeng Chen, Jingyuan Chen, Lin Ma +4
Nowadays, billions of videos are online ready to be viewed and shared. Among an enormous volume of videos, some popular ones are widely viewed by online users while the majority at…
Regularizing RNNs for Caption Generation by Reconstructing The Past with The Present
Xinpeng Chen, Lin Ma, Wenhao Jiang +2
Recently, caption generation with an encoder-decoder framework has been extensively studied and applied in different domains, such as image captioning, code captioning, and so on.…
Learning to Guide Decoding for Image Captioning
Wenhao Jiang, Lin Ma, Xinpeng Chen +2
Recently, much advance has been made in image captioning, and an encoder-decoder framework has achieved outstanding performance for this task. In this paper, we propose an extensio…
Aggregating Frame-level Features for Large-Scale Video Classification
Shaoxiang Chen, Xi Wang, Yongyi Tang +3
This paper introduces the system we developed for the Google Cloud & YouTube-8M Video Understanding Challenge, which can be considered as a multi-label classification problem defin…