33 citations · 57 across the 4 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2022★ 7 cited
Video-Guided Curriculum Learning for Spoken Video Grounding
Yan Xia, Zhou Zhao, Shangwei Ye +3
In this paper, we introduce a new task, spoken video grounding (SVG), which aims to localize the desired video fragments from spoken language descriptions. Compared with using text…
cs.CV2020★ 17 cited
Invariant Deep Compressible Covariance Pooling for Aerial Scene Categorization
Shidong Wang, Yi Ren, Gerard Parr +2
Learning discriminative and invariant feature representation is the key to visual image categorization. In this article, we propose a novel invariant deep compressible covariance p…