14 citations · 25 across the 2 of their papers we have counts for
3 papers
cs.CV2020★ 11 cited
Learning Modality Interaction for Temporal Sentence Localization and Event Captioning in Videos
Shaoxiang Chen, Wenhao Jiang, Wei Liu +1
Automatically generating sentences to describe events and temporally localizing sentences in a video are two important tasks that bridge language and videos. Recent techniques leve…
cs.CV2018
Non-local NetVLAD Encoding for Video Classification
Yongyi Tang, Xing Zhang, Jingwen Wang +3
This paper describes our solution for the 2 YouTube-8M video understanding challenge organized by Google AI. Unlike the video recognition benchmarks, such as Kinetics a…
cs.CV2017★ 14 cited
Aggregating Frame-level Features for Large-Scale Video Classification
Shaoxiang Chen, Xi Wang, Yongyi Tang +3
This paper introduces the system we developed for the Google Cloud & YouTube-8M Video Understanding Challenge, which can be considered as a multi-label classification problem defin…