most citedAggregating Frame-level Features for Large-Scale Video Classification

14 citations · 14 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CV2018

Real-Time Referring Expression Comprehension by Single-Stage Grounding Network

Xinpeng Chen, Lin Ma, Jingyuan Chen +3

In this paper, we propose a novel end-to-end model, namely Single-Stage Grounding network (SSG), to localize the referent given a referring expression within an image. Different fr…

cs.CV2018

Fine-grained Video Attractiveness Prediction Using Multimodal Deep Learning on a Large Real-world Dataset

Xinpeng Chen, Jingyuan Chen, Lin Ma +4

Nowadays, billions of videos are online ready to be viewed and shared. Among an enormous volume of videos, some popular ones are widely viewed by online users while the majority at…

cs.CV2018

Regularizing RNNs for Caption Generation by Reconstructing The Past with The Present

Xinpeng Chen, Lin Ma, Wenhao Jiang +2

Recently, caption generation with an encoder-decoder framework has been extensively studied and applied in different domains, such as image captioning, code captioning, and so on.…

cs.CV2018

Learning to Guide Decoding for Image Captioning

Wenhao Jiang, Lin Ma, Xinpeng Chen +2

Recently, much advance has been made in image captioning, and an encoder-decoder framework has achieved outstanding performance for this task. In this paper, we propose an extensio…

cs.CV201714 cited

Aggregating Frame-level Features for Large-Scale Video Classification

Shaoxiang Chen, Xi Wang, Yongyi Tang +3

This paper introduces the system we developed for the Google Cloud & YouTube-8M Video Understanding Challenge, which can be considered as a multi-label classification problem defin…