195 citations · 210 across the 3 of their papers we have counts for
3 papers
cs.CV2021★ 1 cited
What Makes for Hierarchical Vision Transformer?
Yuxin Fang, Xinggang Wang, Rui Wu +1
Recent studies indicate that hierarchical Vision Transformer with a macro architecture of interleaved non-overlapped window-based self-attention \& shifted-window operation is able…
cs.CV2021★ 195 cited
You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
Yuxin Fang, Bencheng Liao, Xinggang Wang +5
Can Transformer perform 2D object- and region-level recognition from a pure sequence-to-sequence perspective with minimal knowledge about the 2D spatial structure? To answer this q…
cs.CV2019★ 14 cited
Exploiting Offset-guided Network for Pose Estimation and Tracking
Rui Zhang, Zheng Zhu, Peng Li +4
Human pose estimation has witnessed a significant advance thanks to the development of deep learning. Recent human pose estimation approaches tend to directly predict the location…