385 citations · 1.5k across the 44 of their papers we have counts for
Showing 2018Show all
2 papers · 1 filter
cs.CV2018★ 23 cited
StNet: Local and Global Spatial-Temporal Modeling for Action Recognition
Dongliang He, Zhichao Zhou, Chuang Gan +5
Despite the success of deep learning for static image understanding, it remains unclear what are the most effective network architectures for the spatial-temporal modeling in video…
cs.CV2018★ 10 cited
Structured Triplet Learning with POS-tag Guided Attention for Visual Question Answering
Zhe Wang, Xiaoyi Liu, Liangjian Chen +4
Visual question answering (VQA) is of significant interest due to its potential to be a strong test of image understanding systems and to probe the connection between language and…