3 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 1 cited
Revisiting Vision Transformer from the View of Path Ensemble
Shuning Chang, Pichao Wang, Hao Luo +2
Vision Transformers (ViTs) are normally regarded as a stack of transformer layers. In this work, we propose a novel view of ViTs showing that they can be seen as ensemble networks…
cs.CV2023
DOAD: Decoupled One Stage Action Detection Network
Shuning Chang, Pichao Wang, Fan Wang +2
Localizing people and recognizing their actions from videos is a challenging task towards high-level video understanding. Existing methods are mostly two-stage based, with one stag…
cs.CV2023★ 3 cited
Making Vision Transformers Efficient from A Token Sparsification View
Shuning Chang, Pichao Wang, Ming Lin +4
The quadratic computational complexity to the number of tokens limits the practical applications of Vision Transformers (ViTs). Several works propose to prune redundant tokens to a…