16 citations · 25 across the 2 of their papers we have counts for
2 papers
cs.CV2022★ 9 cited
Self-supervised Video Representation Learning with Motion-Aware Masked Autoencoders
Haosen Yang, Deng Huang, Bin Wen +5
Masked autoencoders (MAEs) have emerged recently as art self-supervised spatiotemporal representation learners. Inheriting from the image counterparts, however, existing video MAEs…
cs.CV2021★ 16 cited
Temporal Action Proposal Generation with Transformers
Lining Wang, Haosen Yang, Wenhao Wu +2
Transformer networks are effective at modeling long-range contextual information and have recently demonstrated exemplary performance in the natural language processing domain. Con…