17 citations · 36 across the 8 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2020
BiST: Bi-directional Spatio-Temporal Reasoning for Video-Grounded Dialogues
Hung Le, Doyen Sahoo, Nancy F. Chen +1
Video-grounded dialogues are very challenging due to (i) the complexity of videos which contain both spatial and temporal variations, and (ii) the complexity of user utterances whi…
cs.CV2017★ 17 cited
Truly Multi-modal YouTube-8M Video Classification with Video, Audio, and Text
Zhe Wang, Kingsley Kuan, Mathieu Ravaut +13
The YouTube-8M video classification challenge requires teams to classify 0.7 million videos into one or more of 4,716 classes. In this Kaggle competition, we placed in the top 3% o…